Abstract
Clever CSV is an open-source AI & Machine Learning project. CleverCSV is a Python package for handling messy CSV files. It provides a drop-in replacement for the builtin CSV module with improved dialect detection, and comes with a handy command line application for working with CSV files. CleverCSV is a Python package that aims to solve some of the pain points of CSV files, while maintaining many of the good things. The package automatically detects (with high accuracy) the format (dialect) of CSV files, thus making it easier to simply point to a CSV file and load it, without the need for human inspection. It is built using Python. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
CleverCSV is a Python package that aims to solve some of the pain points of CSV files, while maintaining many of the good things. The package automatically detects (with high accuracy) the format (dialect) of CSV files, thus making it easier to simply point to a CSV file and load it, without the need for human inspection. In the future, we hope to solve some of the other issues of CSV files too.
CleverCSV is based on science. We investigated thousands of real-world CSV files to find a robust way to automatically detect the dialect of a file. This may seem like an easy problem, but to a computer a CSV file is simply a long string, and every dialect will give you some table. In CleverCSV we use a technique based on the patterns of row lengths of the parsed file and the data type of the resulting cells. With our method we achieve 97% accuracy for dialect detection, with a 21% improvement on non-standard (messy) CSV files compared to the Python standard library.
And of course, if you like the package please spread the word! You can do this by Tweeting about it (#CleverCSV) or clicking the ⭐ on GitHub!
2. Objective
CleverCSV is a Python package for handling messy CSV files. It provides a drop-in replacement for the builtin CSV module with improved dialect detection, and comes with a handy command line application for working with CSV files.
This project demonstrates how Python can be applied to a real-world AI & Machine Learning problem.
4. Technology Stack
- detect_dialect:
- read_table:
- read_dataframe:
- read_dicts:
- write_table:
- write_dicts:
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later
- pip / virtualenv for dependencies
- VS Code, PyCharm or Jupyter Notebook
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/alan-turing-institute/CleverCSV.git
cd CleverCSV# Import the package
>>> import clevercsv
# Load the file as a list of rows
# This uses the imdb.csv file in the examples directory
>>> rows = clevercsv.read_table('./imdb.csv')
# Load the file as a Pandas Dataframe
# Note that df = pd.read_csv('./imdb.csv') would fail here
>>> df = clevercsv.read_dataframe('./imdb.csv')
# Use CleverCSV as drop-in replacement for the Python CSV module
# This follows the Sniffer example: https://docs.python.org/3/library/csv.html#csv.Sniffer
# Note that csv.Sniffer would fail here
>>> with open('./imdb.csv', newline='') as csvfile:
... dialect = clevercsv.Sniffer().sniff(csvfile.read())
... csvfile.seek(0)
... reader = clevercsv.reader(csvfile, dialect)
... rows = list(reader)# Install the full version of CleverCSV (this includes the command line interface)
$ pip install clevercsv[full]
# Detect the dialect
$ clevercsv detect ./imdb.csv
Detected: SimpleDialect(',', '', '\\')
# Generate code to import the file
$ clevercsv code ./imdb.csv
import clevercsv
with open("./imdb.csv", "r", newline="", encoding="utf-8") as fp:
reader = clevercsv.reader(fp, delimiter=",", quotechar="", escapechar="\\")
rows = list(reader)
# Explore the CSV file as a Pandas dataframe
$ clevercsv explore -p imdb.csv
Dropping you into an interactive shell.
CleverCSV has loaded the data into the variable: df
>>> df$ pip install clevercsv[full]$ pip install clevercsvFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by alan-turing-institute and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship