A production-ready Retrieval-Augmented Generation (RAG) system that lets you upload any PDF and ask natural-language questions. Every answer is grounded in the document and includes citations (page number + passage) so you can verify the source.
Built as a submission for the Acadza Technologies AI/ML Internship.
POST /upload → ingest research_paper.pdf (3 pages, 12 chunks)
POST /query → "What methodology did the authors use?"
→ "The authors used a mixed-methods approach… [Passage 2, Page 1]"
The Streamlit UI wraps these endpoints in a chat interface.
Large Language Models are trained on general data — they hallucinate when asked about domain-specific documents they have never seen (internal reports, research papers, contracts). Simply pasting an entire PDF into a prompt is expensive and hits context-length limits.
RAG solves this by retrieving only the most relevant passages before the LLM sees them, keeping the prompt short and the answer grounded in evidence.
Documents are split with RecursiveCharacterTextSplitter (LangChain):
- Chunk size: 1000 chars, overlap: 200 chars
- The splitter tries paragraph → sentence → word boundaries, in that order, so chunks rarely split mid-sentence.
- Overlap prevents an answer from being "cut" at a chunk boundary.
sentence-transformers/all-MiniLM-L6-v2:
- Free, runs locally — no API costs for embeddings.
- 384-dimensional vectors; strong retrieval quality on English text.
- The same model is used at both ingest and query time so the vector space is consistent.
Maximal Marginal Relevance (MMR) instead of plain top-k cosine similarity:
- Fetches
top_k × 3candidates, then re-ranks to maximise both relevance and diversity. - Avoids returning 5 nearly-identical chunks from the same paragraph.
ChromaDB with local SQLite persistence:
- Ships as a Python package — zero infrastructure to stand up.
- Supports named collections, so multiple PDFs can be kept separate.
- Easy to swap for a cloud vector store (Pinecone, Weaviate) by changing
one constructor call in
ingest.py.
Configurable via LLM_PROVIDER env var:
- OpenAI
gpt-4o-mini(default) — fast, cheap, good instruction-following. - Anthropic Claude Haiku — swap with
LLM_PROVIDER=anthropic.
The prompt (in prompts.py) instructs the model to cite passages explicitly
and to say "I could not find an answer" when context is insufficient —
avoiding confident hallucinations.
| Iteration | What I tried | Result | What changed |
|---|---|---|---|
| v1 | Chunk size 500, no overlap | Answers cut off mid-thought | → increased to 1000 / 200 overlap |
| v2 | Plain cosine top-k retrieval | Redundant passages from same para | → switched to MMR |
| v3 | System prompt without citation instruction | Model mixed doc facts with training data | → added strict "ONLY use context" rule + citation requirement |
| v4 | Single global collection | Multiple PDFs interfered | → added per-request collection param |
Why ChromaDB over Pinecone? Pinecone requires account setup and network calls; ChromaDB runs in-process with zero config. For an internship demo, removing infrastructure friction lets reviewers run it in 2 commands.
Why sentence-transformers over OpenAI embeddings? Removes a second API
dependency. For production with large corpora, switching to text-embedding-3-small
would improve retrieval quality and the code change is one line.
Why MMR? A naive top-5 similarity search on a 10-page PDF often returns 5 sentences from the same dense paragraph. MMR trades a small relevance loss for much better coverage of the document.
Why a separate prompts.py? Prompt engineering is iterative. Keeping
prompts isolated from retrieval logic means I can A/B test wording without
touching the API or database code.
Built over 4 days (~3 hours/day):
| Day | Work done |
|---|---|
| 1 | Project scaffold, FastAPI endpoints, PyPDF ingestion |
| 2 | ChromaDB integration, embedding pipeline, ingest tests |
| 3 | Query pipeline, MMR retrieval, prompt engineering, LLM integration |
| 4 | Streamlit UI, full test suite, README, cleanup |
pdf-rag-assistant/
├── app/
│ ├── main.py # FastAPI app (upload + query endpoints)
│ ├── ingest.py # PDF → chunks → embeddings → ChromaDB
│ ├── query.py # question → retrieve → LLM → answer
│ └── prompts.py # prompt templates (isolated for easy iteration)
├── ui/
│ └── streamlit_app.py # chat interface
├── tests/
│ └── test_query.py # 10 tests covering prompts, API, pipeline
├── .env.example
├── requirements.txt
└── README.md
# 1. Clone and install
git clone https://github.com/yourusername/pdf-rag-assistant.git
cd pdf-rag-assistant
pip install -r requirements.txt
# 2. Set your API key
cp .env.example .env
# Edit .env → add OPENAI_API_KEY (or ANTHROPIC_API_KEY)
# 3. Start the API server
uvicorn app.main:app --reload --port 8000
# 4. Start the UI (in a separate terminal)
streamlit run ui/streamlit_app.py
# 5. Run tests
pytest tests/ -vOpen http://localhost:8501 → upload a PDF → start asking questions.
| Parameter | Type | Default | Description |
|---|---|---|---|
file |
required | PDF to ingest | |
collection |
string | "default" |
Named collection (separate PDFs) |
chunk_size |
int | 1000 |
Characters per chunk |
chunk_overlap |
int | 200 |
Overlap between chunks |
{
"question": "What are the main findings?",
"top_k": 5,
"collection": "default"
}Response:
{
"answer": "The main findings are… [Passage 1, Passage 3]",
"sources": [
{"text": "...", "page": 2, "source_file": "paper.pdf"}
],
"collection": "default"
}Python · FastAPI · LangChain · ChromaDB · sentence-transformers · OpenAI / Anthropic · Streamlit · pytest