Fake News Detector

A professional TF-IDF + Logistic Regression style-risk classifier for educational fake-news detection, with a Streamlit dashboard, honest evaluation, uncertainty handling, and leakage analysis.

AI & Machine LearningPythonMIT

Abstract

Fake News Detector is an open-source AI & Machine Learning project. A professional TF-IDF + Logistic Regression style-risk classifier for educational fake-news detection, with a Streamlit dashboard, honest evaluation, uncertainty handling, and leakage analysis. This project demonstrates an end-to-end, honest text-classification workflow on a labeled news dataset. It includes preprocessing, model training, threshold-based decisions with an uncertainty band, dataset leakage analysis, leakage-controlled training, source-confounding diagnostics, visual reports, checksum-verified artifacts, and a Streamlit dashboard. It is built using Python, Machine Learning, NLP. Key capabilities include: TF-IDF vectorization with word n-grams for text feature extraction; Logistic Regression baseline classifier; Pipeline-embedded TextCleaner so training and inference clean text identically (no train/serve skew). The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.

1. Introduction

This project demonstrates an end-to-end, honest text-classification workflow on a labeled news dataset. It includes preprocessing, model training, threshold-based decisions with an uncertainty band, dataset leakage analysis, leakage-controlled training, source-confounding diagnostics, visual reports, checksum-verified artifacts, and a Streamlit dashboard.

The goal is to show how a text classifier can be turned into a responsible decision-support tool, not just a single accuracy or AUC score.

A responsible machine learning project that turns news text into a style-risk signal, using a TF-IDF + Logistic Regression pipeline with honest evaluation, dataset leakage analysis, leakage-controlled training, checksum-verified model loading, a Streamlit dashboard, and command-line inference.

2. Objective

A professional TF-IDF + Logistic Regression style-risk classifier for educational fake-news detection, with a Streamlit dashboard, honest evaluation, uncertainty handling, and leakage analysis.

This project demonstrates how Python, Machine Learning, NLP can be applied to a real-world AI & Machine Learning problem.

3. Key Features / Modules

  • TF-IDF vectorization with word n-grams for text feature extraction
  • Logistic Regression baseline classifier
  • Pipeline-embedded TextCleaner so training and inference clean text identically (no train/serve skew)
  • REAL / FAKE / UNCERTAIN output with a configurable uncertainty band
  • Dataset leakage report with a "contains Reuters" heuristic
  • Source-confounding diagnostic with a quantified confounding score
  • Leakage-controlled training via source-artifact stripping
  • Out-of-source evaluation that runs when the data permits and reports infeasibility when it does not
  • Checksum-verified model loading with SHA-256 sidecars
  • Streamlit dashboard for interactive style-risk analysis

4. Technology Stack

PythonMachine LearningNLP
  • scikit-learn
  • matplotlib
  • Streamlit
  • GitHub Actions

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/AmirhosseinHonardoust/Fake-News-Detector.git
cd Fake-News-Detector
pip install -r requirements.txt
pip install -r requirements-dev.txt
python src/train_model.py
streamlit run src/streamlit_app.py

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Deploy the model as a web app with Streamlit, Flask or FastAPI
  • Compare against an additional model and report the metric difference
  • Add explainability (SHAP / Grad-CAM)

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What dataset does the project use and how was it pre-processed?
  2. Which algorithm / model architecture is used and why was it chosen over alternatives?
  3. How are training and testing data split, and how is overfitting avoided?
  4. Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
  5. How would you deploy this model for real users?

9. Source Code & License

This project is developed by AmirhosseinHonardoust and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for AI & Machine Learning Internship