azure-doc-rag-citations is a local-first Retrieval-Augmented Generation (RAG) repository built for production-style engineering workflows.
It ingests local .txt and .md files, stores deterministic offline embeddings in Qdrant, and serves grounded answers through a FastAPI endpoint with chunk-level citations.
Why this matters in production RAG:
- Deterministic embeddings and extractive generation improve auditability and reproducibility.
- Citation-first responses make answers inspectable.
- Local mode is fully offline and requires zero cloud credentials.
- Azure adapters and IaC skeleton are included for future cloud deployment.
- Local-first and offline by default (no Azure subscription required)
- Deterministic embeddings via
HashingVectorizer(n_features=1024,alternate_sign=False,norm=None) - Qdrant local vector store via Docker Compose
- Grounded extractive generation that only uses retrieved text
- Citation-rich API responses (
doc_id,title,chunk_id,score,snippet) - Structured JSON logging with
request_idand latency - Optional Azure adapters (Azure OpenAI + Azure AI Search stub)
- Windows-first PowerShell scripts for dev/test/demo
flowchart LR
A[Local Docs .md/.txt] --> B[Ingestion CLI]
B --> C[Chunking Engine]
C --> D[HashingVectorizer Embeddings]
D --> E[Qdrant Local Vector DB]
F[FastAPI /chat] --> G[RAG Service]
G --> H[Retrieval Provider]
H --> E
G --> I[Generation Provider]
I --> J[Local Extractive Generator]
I --> K[Optional Azure OpenAI Generator]
G --> L[Answer + Citations + Metadata]
# 1) Start local stack + API server
.\scripts\dev.ps1In another PowerShell window:
# 2) Ingest sample docs
.\.venv\Scripts\python.exe -m app.ingest --path data/sample_docs --collection docs
# 3) Ask a question
Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/chat -ContentType "application/json" -Body (@{
question = "What should be attached before production deployment?"
top_k = 5
mode = "local"
} | ConvertTo-Json).\scripts\demo.ps1 -Question "What should be attached before production deployment?"docker compose up -d
python -m venv .venv
source .venv/bin/activate
pip install -e .
python -m app.ingest --path data/sample_docs --collection docs
uvicorn app.api:app --host 127.0.0.1 --port 8000curl http://127.0.0.1:8000/healthcurl -X POST http://127.0.0.1:8000/chat \
-H "Content-Type: application/json" \
-d '{
"question": "Who approves emergency changes?",
"top_k": 5,
"mode": "local"
}'$body = @{
question = "Who approves emergency changes?"
top_k = 5
mode = "local"
} | ConvertTo-Json
Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/chat -ContentType "application/json" -Body $bodyEvaluation input lives at data/eval/questions.json.
Run:
.\.venv\Scripts\python.exe -m app.eval --questions data/eval/questions.json --collection docsThe script prints recall@1, recall@3, and recall@5.
What recall@k means:
- For each question, check whether expected
doc_ids are in the top-k retrieved results. - Average that fraction across all evaluation questions.
Local mode does not require Azure.
Optional adapters are provided:
src/app/providers/generation_azure_openai.pysrc/app/providers/retrieval_azure_ai_search.py(stub)infra/azure/IaC skeleton
To enable Azure OpenAI mode, set:
AZURE_OPENAI_ENDPOINTAZURE_OPENAI_API_KEYAZURE_OPENAI_DEPLOYMENT
To implement Azure AI Search retrieval, set:
AZURE_AI_SEARCH_ENDPOINTAZURE_AI_SEARCH_API_KEYAZURE_AI_SEARCH_INDEX
HashingVectorizervs semantic embedding models:- Pros: deterministic, fast, no model downloads, no internet required.
- Cons: weaker semantic recall than transformer embeddings.
- Local extractive generation vs generative LLM:
- Pros: grounded-by-construction, citation-safe, no hallucinated synthesis.
- Cons: less fluent and less abstractive than LLM outputs.
- No secrets are committed.
- Use
.env(not tracked) for credentials. mode=localperforms no outbound Azure calls.- Review
SECURITY.mdfor disclosure guidance.
# lint + tests
.\scripts\test.ps1
# run demo end-to-end
.\scripts\demo.ps1MIT (LICENSE).