Abstract
Resume Screening RAG Pipeline is an open-source AI & Machine Learning project. An LLM Chatbot that dynamically retrieves and processes resumes using RAG to perform resume screening. The research is part of the author's graduating thesis, which aims to present a POC of an LLM chatbot that can assist hiring managers in the resume screening process. The assistant is a cost-efficient, user-friendly, and more effective alternative to the conventional keyword-based screening methods. It is built using Jupyter Notebook, LangChain, OpenAI API. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
The research is part of the author's graduating thesis, which aims to present a POC of an LLM chatbot that can assist hiring managers in the resume screening process. The assistant is a cost-efficient, user-friendly, and more effective alternative to the conventional keyword-based screening methods. Powered by state-of-the-art LLMs, it can handle unstructured and complex natural language data in job descriptions/resumes while performing high-level tasks as effectively as a human recruiter.
Despite the increasingly large volume of applicants each year, there are limited tools that can assist the screening process effectively and reliably. Existing methods often revolve around keyword-based approaches, which cannot accurately handle the complexity of natural language in human-written documents. Because of this, there is a clear opportunity to integrate LLM-based methods into this domain, which the project aims to address.
2. Objective
An LLM Chatbot that dynamically retrieves and processes resumes using RAG to perform resume screening.
This project demonstrates how Jupyter Notebook, LangChain, OpenAI API can be applied to a real-world AI & Machine Learning problem.
4. Technology Stack
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later with Jupyter Notebook / JupyterLab (or Google Colab)
- pip for dependencies
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/Hungreeee/Resume-Screening-RAG-Pipeline.git
cd Resume-Screening-RAG-Pipeline# Clone the project
git clone https://github.com/Hungreeee/Resume-Screening-RAG-Pipeline.git
# Install dependencies
pip install requirements.txtstreamlit run demo/interface.pyFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by Hungreeee and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship