Docmind AI LLM

DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline.

Data SciencePythonMIT

Abstract

Docmind AI LLM is an open-source Data Science project. DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline. semantic_search plus configured hybrid_search, keyword_search, multimodal_search, and knowledge_graph tools. It is built using Python, LangChain. Key capabilities include: Privacy-focused, local-first: Remote LLM endpoints are blocked by default; enable explicitly when needed; Multi-format parsing: Docling covers PDFs and common office/HTML formats; only explicit text formats use the direct UTF-8 loader. Binary parser failures stop ingestion instead of decoding source bytes as text; Hybrid retrieval with routing: RouterQueryEngine with required. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

semantic_search plus configured hybrid_search, keyword_search, multimodal_search, and knowledge_graph tools.

remote endpoints must be explicitly enabled. The release benchmark does not claim independently measured zero egress.

2. Objective

DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline.

This project demonstrates how Python, LangChain can be applied to a real-world Data Science problem.

3. Key Features / Modules

  • Privacy-focused, local-first: Remote LLM endpoints are blocked by default; enable explicitly when needed.
  • Multi-format parsing: Docling covers PDFs and common office/HTML formats; only explicit text formats use the direct UTF-8 loader. Binary parser failures stop ingestion instead of decoding source bytes as text.
  • Hybrid retrieval with routing: RouterQueryEngine with required
  • Qdrant server-side fusion: Query API RRF (default) or DBSF over named vectors text-dense and text-sparse; sparse queries use FastEmbed BM42/BM25 when available.
  • Reranking and multimodal: Text rerank uses a BGE cross-encoder; SigLIP reranks visual nodes.
  • Multi-agent coordination: LangGraph supervisor orchestrates four agents (planner, retrieval, synthesis, validation); LlamaIndex owns retrieval routing.
  • Snapshots and reproducibility: Qdrant owns live vectors; app snapshots bind
  • PDF page images: pypdfium2 renders page images to WebP/JPEG; optional AES-GCM encryption with .enc outputs and just-in-time decryption for visual scoring.
  • ArtifactStore (multimodal durability): Page images/thumbnails are stored as content-addressed ArtifactRef(sha256, suffix) (no base64 blobs or host paths in durable stores).
  • Multimodal UX: Chat renders image sources and supports query-by-image “Visual search” (SigLIP) for image-rich PDFs.

4. Technology Stack

PythonLangChain

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/BjornMelin/docmind-ai-llm.git
cd docmind-ai-llm
  1. Clone the repository:
  2. Install dependencies:
  3. LlamaIndex Core (>=0.14.21,<0.15.0): Ingestion, retrieval, selectors, and query engines, with selected LLM, Hugging Face, Qdrant, and DuckDB adapters
  4. LangGraph (>=1.0.10,<2.0.0): Four-worker supervisor orchestration (graph-native StateGraph, no external supervisor wrapper)
  5. Streamlit (>=1.52.2,<2.0.0): Web interface framework
  6. Ollama (0.6.2): Local LLM integration
  7. Qdrant Client (>=1.15.1,<2.0.0): Vector database operations
  8. Docling (>=2.111,<3): Multi-format document conversion.
git clone https://github.com/BjornMelin/docmind-ai-llm.git
   cd docmind-ai-llm
uv sync --frozen
uv sync --frozen --extra observability
uv sync --frozen --extra searchable-pdf

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by BjornMelin and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship