Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PDF RAG Assistant

A production-ready Retrieval-Augmented Generation (RAG) system that lets you upload any PDF and ask natural-language questions. Every answer is grounded in the document and includes citations (page number + passage) so you can verify the source.

Built as a submission for the Acadza Technologies AI/ML Internship.


Demo

POST /upload  →  ingest research_paper.pdf  (3 pages, 12 chunks)
POST /query   →  "What methodology did the authors use?"
               →  "The authors used a mixed-methods approach… [Passage 2, Page 1]"

The Streamlit UI wraps these endpoints in a chat interface.


Problem

Large Language Models are trained on general data — they hallucinate when asked about domain-specific documents they have never seen (internal reports, research papers, contracts). Simply pasting an entire PDF into a prompt is expensive and hits context-length limits.

RAG solves this by retrieving only the most relevant passages before the LLM sees them, keeping the prompt short and the answer grounded in evidence.


Approach

Chunking strategy

Documents are split with RecursiveCharacterTextSplitter (LangChain):

  • Chunk size: 1000 chars, overlap: 200 chars
  • The splitter tries paragraph → sentence → word boundaries, in that order, so chunks rarely split mid-sentence.
  • Overlap prevents an answer from being "cut" at a chunk boundary.

Embedding model

sentence-transformers/all-MiniLM-L6-v2:

  • Free, runs locally — no API costs for embeddings.
  • 384-dimensional vectors; strong retrieval quality on English text.
  • The same model is used at both ingest and query time so the vector space is consistent.

Retrieval

Maximal Marginal Relevance (MMR) instead of plain top-k cosine similarity:

  • Fetches top_k × 3 candidates, then re-ranks to maximise both relevance and diversity.
  • Avoids returning 5 nearly-identical chunks from the same paragraph.

Storage

ChromaDB with local SQLite persistence:

  • Ships as a Python package — zero infrastructure to stand up.
  • Supports named collections, so multiple PDFs can be kept separate.
  • Easy to swap for a cloud vector store (Pinecone, Weaviate) by changing one constructor call in ingest.py.

LLM

Configurable via LLM_PROVIDER env var:

  • OpenAI gpt-4o-mini (default) — fast, cheap, good instruction-following.
  • Anthropic Claude Haiku — swap with LLM_PROVIDER=anthropic.

The prompt (in prompts.py) instructs the model to cite passages explicitly and to say "I could not find an answer" when context is insufficient — avoiding confident hallucinations.


Iterations

Iteration What I tried Result What changed
v1 Chunk size 500, no overlap Answers cut off mid-thought → increased to 1000 / 200 overlap
v2 Plain cosine top-k retrieval Redundant passages from same para → switched to MMR
v3 System prompt without citation instruction Model mixed doc facts with training data → added strict "ONLY use context" rule + citation requirement
v4 Single global collection Multiple PDFs interfered → added per-request collection param

Key design choices

Why ChromaDB over Pinecone? Pinecone requires account setup and network calls; ChromaDB runs in-process with zero config. For an internship demo, removing infrastructure friction lets reviewers run it in 2 commands.

Why sentence-transformers over OpenAI embeddings? Removes a second API dependency. For production with large corpora, switching to text-embedding-3-small would improve retrieval quality and the code change is one line.

Why MMR? A naive top-5 similarity search on a 10-page PDF often returns 5 sentences from the same dense paragraph. MMR trades a small relevance loss for much better coverage of the document.

Why a separate prompts.py? Prompt engineering is iterative. Keeping prompts isolated from retrieval logic means I can A/B test wording without touching the API or database code.


Daily time commitment

Built over 4 days (~3 hours/day):

Day Work done
1 Project scaffold, FastAPI endpoints, PyPDF ingestion
2 ChromaDB integration, embedding pipeline, ingest tests
3 Query pipeline, MMR retrieval, prompt engineering, LLM integration
4 Streamlit UI, full test suite, README, cleanup

Project structure

pdf-rag-assistant/
├── app/
│   ├── main.py        # FastAPI app (upload + query endpoints)
│   ├── ingest.py      # PDF → chunks → embeddings → ChromaDB
│   ├── query.py       # question → retrieve → LLM → answer
│   └── prompts.py     # prompt templates (isolated for easy iteration)
├── ui/
│   └── streamlit_app.py   # chat interface
├── tests/
│   └── test_query.py      # 10 tests covering prompts, API, pipeline
├── .env.example
├── requirements.txt
└── README.md

Quick start

# 1. Clone and install
git clone https://github.com/yourusername/pdf-rag-assistant.git
cd pdf-rag-assistant
pip install -r requirements.txt

# 2. Set your API key
cp .env.example .env
# Edit .env → add OPENAI_API_KEY (or ANTHROPIC_API_KEY)

# 3. Start the API server
uvicorn app.main:app --reload --port 8000

# 4. Start the UI (in a separate terminal)
streamlit run ui/streamlit_app.py

# 5. Run tests
pytest tests/ -v

Open http://localhost:8501 → upload a PDF → start asking questions.


API reference

POST /upload

Parameter Type Default Description
file PDF required PDF to ingest
collection string "default" Named collection (separate PDFs)
chunk_size int 1000 Characters per chunk
chunk_overlap int 200 Overlap between chunks

POST /query

{
  "question": "What are the main findings?",
  "top_k": 5,
  "collection": "default"
}

Response:

{
  "answer": "The main findings are… [Passage 1, Passage 3]",
  "sources": [
    {"text": "...", "page": 2, "source_file": "paper.pdf"}
  ],
  "collection": "default"
}

Stack

Python · FastAPI · LangChain · ChromaDB · sentence-transformers · OpenAI / Anthropic · Streamlit · pytest

About

RAG-powered PDF Q&A — upload any PDF, ask questions, get cited answers grounded in the document. Built with FastAPI, LangChain, ChromaDB & LLMs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages