Abstract
Benchm ML is an open-source AI & Machine Learning project. A minimal benchmark for scalability, speed and accuracy of commonly used open source implementations (R packages, Python scikit-learn, H2O, xgboost, Spark MLlib etc.) of the top machine learning algorithms for binary classification (random forests, gradient boosted trees, deep neural networks etc.). This project aims at a minimal benchmark for scalability, speed and accuracy of commonly used implementations of a few machine learning algorithms. The target of this study is binary classification with numeric and categorical inputs (of limited cardinality i.e. It is built using R, Machine Learning, Python, Deep Learning. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
This project aims at a minimal benchmark for scalability, speed and accuracy of commonly used implementations of a few machine learning algorithms. The target of this study is binary classification with numeric and categorical inputs (of limited cardinality i.e. not very sparse) and no missing data, perhaps the most common problem in business applications (e.g. credit scoring, fraud detection or churn prediction). If the input matrix is of n x p, n is varied as 10K, 100K, 1M, 10M, while p is ~1K (after expanding the categoricals into dummy variables/one-hot encoding). This particular type of data structure/size (the largest) stems from this author's interest in some particular business applications.
at the end of this repo a summary of how the focus has changed over time, and why instead of updating this benchmark I started a new one (and where to find it).
in various commonly used open source implementations like
2. Objective
A minimal benchmark for scalability, speed and accuracy of commonly used open source implementations (R packages, Python scikit-learn, H2O, xgboost, Spark MLlib etc.) of the top machine learning algorithms for binary classification (random forests, gradient boosted trees, deep neural networks etc.).
This project demonstrates how R, Machine Learning, Python can be applied to a real-world AI & Machine Learning problem.
4. Technology Stack
- linear (logistic regression, linear SVM)
- random forest
- boosting
- deep neural network
- R packages
- Python scikit-learn
- Vowpal Wabbit
- lightgbm (added in 2017)
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later
- pip / virtualenv for dependencies
- VS Code, PyCharm or Jupyter Notebook
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/szilard/benchm-ml.git
cd benchm-mlFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by szilard and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship