Cleanvision

Automatically find issues in image datasets and practice data-centric computer vision.

Data SciencePythonApache-2.0

Abstract

Cleanvision is an open-source Data Science project. Automatically find issues in image datasets and practice data-centric computer vision. CleanVision automatically detects potential issues in image datasets like images that are: blurry, under/over-exposed, (near) duplicates, etc. This data-centric AI package is a quick first step for any computer vision project to find problems in the dataset, which you want to address before applying machine learning. It is built using Python, Computer Vision, Deep Learning. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

CleanVision automatically detects potential issues in image datasets like images that are: blurry, under/over-exposed, (near) duplicates, etc. This data-centric AI package is a quick first step for any computer vision project to find problems in the dataset, which you want to address before applying machine learning. CleanVision is super simple -- run the same couple lines of Python code to audit any image dataset!

2. Objective

Automatically find issues in image datasets and practice data-centric computer vision.

This project demonstrates how Python, Computer Vision, Deep Learning can be applied to a real-world Data Science problem.

4. Technology Stack

PythonComputer VisionDeep Learning

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/cleanlab/cleanvision.git
cd cleanvision
  1. Run CleanVision to audit the images.
  2. CleanVision diagnoses many types of issues, but you can also check for only specific issues.
pip install cleanvision
wget -nc 'https://cleanlab-public.s3.amazonaws.com/CleanVision/image_files.zip'
from cleanvision import Imagelab

# Specify path to folder containing the image files in your dataset
imagelab = Imagelab(data_path="FOLDER_WITH_IMAGES/")

# Automatically check for a predefined list of issues within your dataset
imagelab.find_issues()

# Produce a neat report of the issues found in your dataset
imagelab.report()
issue_types = {"dark": {}, "blurry": {}}

imagelab.find_issues(issue_types=issue_types)

# Produce a report with only the specified issue_types
imagelab.report(issue_types=issue_types)

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by cleanlab and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship