Aio2025 Mlops Project01

MLOps End-to-End: Customer Churn Prediction System

Data ScienceHTMLApache-2.0

Abstract

Aio2025 Mlops Project01 is an open-source Data Science project. MLOps End-to-End: Customer Churn Prediction System. The system is designed for scalability, reproducibility, and production deployment. It is built using HTML. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

The system is designed for scalability, reproducibility, and production deployment.

A production-ready MLOps system for customer churn prediction, demonstrating best practices in machine learning operations including data versioning, feature stores, experiment tracking, model serving, and infrastructure automation.

2. Objective

MLOps End-to-End: Customer Churn Prediction System

This project demonstrates how HTML can be applied to a real-world Data Science problem.

4. Technology Stack

HTML
  • Docker: 20.10+
  • Docker Compose: v2.0+
  • Kubernetes (optional): kubectl + cluster (minikube/kind/cloud)
  • Git: For repository management
  • CUDA (optional): For GPU-accelerated training

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • A modern web browser
  • VS Code or any code editor
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/ThuanNaN/aio2025-mlops-project01.git
cd aio2025-mlops-project01
  1. MLflow: Tracking server + PostgreSQL + MinIO
  2. Kafka: 3-node cluster (KRaft mode) + Kafka UI
  3. Airflow: Airflow 3.x with scheduler, webserver, PostgreSQL
  4. Monitoring: Prometheus + Grafana + Loki
  5. Namespace: mlops
  6. Services: PostgreSQL, MinIO, MLflow, Kafka (3-node), Airflow 3.x
  7. Storage: PersistentVolumeClaims for data persistence
  8. Dashboard: Kubernetes Dashboard for cluster management
cd data-pipeline

# Create virtual environment
conda create -n churn_mlops python=3.10 -y
conda activate churn_mlops

# Install dependencies
pip install -r requirements.txt

# Pull versioned data
dvc pull

# Initialize Feast
cd churn_feature_store
feast apply
cd infra/docker

# Start all services
./run.sh start all

# Start specific service
./run.sh start mlflow

# Stop services
./run.sh stop all

# Check status
./run.sh status
cd infra/k8s

# Deploy all resources
./deploy.sh

# Check status
kubectl get all -n mlops

# Get dashboard token
./get-dashboard-token.sh
kubectl port-forward -n mlops svc/mlflow 5000:5000
kubectl port-forward -n mlops svc/airflow-webserver 8080:8080
kubectl port-forward -n mlops svc/minio-console 9001:9001

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by ThuanNaN and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship