Abstract
Scarff is an open-source AI & Machine Learning project. SCARFF (SCAlable Real-time Frauds Finder) is a framework which enables credit card fraud detection. SCARFF (SCAlable Real-time Fraud Finder) is a framework which enables credit card fraud detection. It is built using Scala. The complete source code is publicly available on GitHub under the GNU General Public License v3.0, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
SCARFF (SCAlable Real-time Fraud Finder) is a framework which enables credit card fraud detection.
SCAlable Real-time Fraud Finder (SCARFF) is an open source platform which processes and analyses credit card streaming data in order to return reliable alerts in a nearly real-time setting. This original framework for near real-time Streaming Fraud Detection integrates Big Data tools (Kafka, Spark and Cassandra) with a machine learning approach which deals with data imbalance, non-stationarity and feedback latency.
At the core of SCARFF there is a Spark application and here we present its implementation.
2. Objective
SCARFF (SCAlable Real-time Frauds Finder) is a framework which enables credit card fraud detection.
This project demonstrates how Scala can be applied to a real-world AI & Machine Learning problem.
4. Technology Stack
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- See the project README for exact requirements
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/fabriziocarcillo/scarff.git
cd scarffFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by fabriziocarcillo and published on GitHub under the GNU General Public License v3.0. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship