Abstract
Movie Recommendation System is an open-source AI & Machine Learning project. Movie recommendation system with Python. Implements content-based filtering (TF-IDF + cosine similarity), collaborative filtering with matrix factorization (TruncatedSVD), and a hybrid approach. Evaluates with Precision@K, Recall@K, and NDCG. Includes rating distribution plots, top movies, and sample recommendations. This project demonstrates an end-to-end movie recommendation workflow using synthetic movie metadata and rating interactions. It includes synthetic data generation, user-level train/test splitting, content-based recommendations, collaborative recommendations, hybrid recommendations, baseline comparisons, alpha-sweep evaluation, visual reports, structured recommendation exports, automated tests, and GitHub Actions CI. It is built using Python, Machine Learning. Key capabilities include: Synthetic movie and ratings generator with deterministic seed behavior; Content-based filtering using TF-IDF over movie genres; Collaborative filtering using a sparse user-item matrix and TruncatedSVD. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
This project demonstrates an end-to-end movie recommendation workflow using synthetic movie metadata and rating interactions. It includes synthetic data generation, user-level train/test splitting, content-based recommendations, collaborative recommendations, hybrid recommendations, baseline comparisons, alpha-sweep evaluation, visual reports, structured recommendation exports, automated tests, and GitHub Actions CI.
A production-minded movie recommendation workflow for comparing content-based filtering, collaborative filtering, hybrid ranking, simple baselines, alpha-sweep evaluation, and structured recommendation outputs on a synthetic ratings dataset.
The goal is to show how a recommender-system demo can be evaluated honestly, not just how to generate recommendations.
2. Objective
Movie recommendation system with Python. Implements content-based filtering (TF-IDF + cosine similarity), collaborative filtering with matrix factorization (TruncatedSVD), and a hybrid approach. Evaluates with Precision@K, Recall@K, and NDCG. Includes rating distribution plots, top movies, and sample recommendations.
This project demonstrates how Python, Machine Learning can be applied to a real-world AI & Machine Learning problem.
3. Key Features / Modules
- Synthetic movie and ratings generator with deterministic seed behavior
- Content-based filtering using TF-IDF over movie genres
- Collaborative filtering using a sparse user-item matrix and TruncatedSVD
- Corrected SVD reconstruction using scikit-learn's inverse_transform
- Hybrid recommender that blends content and collaborative scores
- Alpha sweep to evaluate the hybrid blend from content-only to collaborative-only
- Baseline comparison against random, popularity, average-rating, Bayesian-average, and positive-count recommenders
- Ranking metrics with Precision@K, Recall@K, and NDCG@K
- Structured recommendation exports with rank, movie ID, title, genres, score, and reason
- Readable text recommendation files for sample users
4. Technology Stack
- scikit-learn
- matplotlib
- unittest
- ruff, black, mypy
- GitHub Actions
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later
- pip / virtualenv for dependencies
- VS Code, PyCharm or Jupyter Notebook
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/AmirhosseinHonardoust/Movie-Recommendation-System.git
cd Movie-Recommendation-Systempip install -r requirements.txtpython data/generate_ratings.py --users 800 --movies 1200 --seed 42 --outdir datapython src/build_recommender.py --ratings data/ratings.csv --movies data/movies.csv --outdir outputs --k 10 --alpha 0.6 --seed 42python -m unittest discover -s tests -vFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by AmirhosseinHonardoust and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship