Skip to content

Repository files navigation

azure-doc-rag-citations

Overview

azure-doc-rag-citations is a local-first Retrieval-Augmented Generation (RAG) repository built for production-style engineering workflows.

It ingests local .txt and .md files, stores deterministic offline embeddings in Qdrant, and serves grounded answers through a FastAPI endpoint with chunk-level citations.

Why this matters in production RAG:

  • Deterministic embeddings and extractive generation improve auditability and reproducibility.
  • Citation-first responses make answers inspectable.
  • Local mode is fully offline and requires zero cloud credentials.
  • Azure adapters and IaC skeleton are included for future cloud deployment.

Features

  • Local-first and offline by default (no Azure subscription required)
  • Deterministic embeddings via HashingVectorizer (n_features=1024, alternate_sign=False, norm=None)
  • Qdrant local vector store via Docker Compose
  • Grounded extractive generation that only uses retrieved text
  • Citation-rich API responses (doc_id, title, chunk_id, score, snippet)
  • Structured JSON logging with request_id and latency
  • Optional Azure adapters (Azure OpenAI + Azure AI Search stub)
  • Windows-first PowerShell scripts for dev/test/demo

Architecture

flowchart LR
    A[Local Docs .md/.txt] --> B[Ingestion CLI]
    B --> C[Chunking Engine]
    C --> D[HashingVectorizer Embeddings]
    D --> E[Qdrant Local Vector DB]

    F[FastAPI /chat] --> G[RAG Service]
    G --> H[Retrieval Provider]
    H --> E
    G --> I[Generation Provider]
    I --> J[Local Extractive Generator]
    I --> K[Optional Azure OpenAI Generator]
    G --> L[Answer + Citations + Metadata]
Loading

Quickstart

Windows PowerShell (recommended)

# 1) Start local stack + API server
.\scripts\dev.ps1

In another PowerShell window:

# 2) Ingest sample docs
.\.venv\Scripts\python.exe -m app.ingest --path data/sample_docs --collection docs

# 3) Ask a question
Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/chat -ContentType "application/json" -Body (@{
  question = "What should be attached before production deployment?"
  top_k = 5
  mode = "local"
} | ConvertTo-Json)

End-to-end demo script

.\scripts\demo.ps1 -Question "What should be attached before production deployment?"

Linux/macOS equivalent (if needed)

docker compose up -d
python -m venv .venv
source .venv/bin/activate
pip install -e .
python -m app.ingest --path data/sample_docs --collection docs
uvicorn app.api:app --host 127.0.0.1 --port 8000

API Usage Examples

Health check

curl http://127.0.0.1:8000/health

Chat with curl

curl -X POST http://127.0.0.1:8000/chat \
  -H "Content-Type: application/json" \
  -d '{
    "question": "Who approves emergency changes?",
    "top_k": 5,
    "mode": "local"
  }'

Chat with PowerShell

$body = @{
  question = "Who approves emergency changes?"
  top_k = 5
  mode = "local"
} | ConvertTo-Json

Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/chat -ContentType "application/json" -Body $body

Evaluation

Evaluation input lives at data/eval/questions.json.

Run:

.\.venv\Scripts\python.exe -m app.eval --questions data/eval/questions.json --collection docs

The script prints recall@1, recall@3, and recall@5.

What recall@k means:

  • For each question, check whether expected doc_ids are in the top-k retrieved results.
  • Average that fraction across all evaluation questions.

Azure-ready

Local mode does not require Azure.

Optional adapters are provided:

  • src/app/providers/generation_azure_openai.py
  • src/app/providers/retrieval_azure_ai_search.py (stub)
  • infra/azure/ IaC skeleton

To enable Azure OpenAI mode, set:

  • AZURE_OPENAI_ENDPOINT
  • AZURE_OPENAI_API_KEY
  • AZURE_OPENAI_DEPLOYMENT

To implement Azure AI Search retrieval, set:

  • AZURE_AI_SEARCH_ENDPOINT
  • AZURE_AI_SEARCH_API_KEY
  • AZURE_AI_SEARCH_INDEX

Tradeoffs

  • HashingVectorizer vs semantic embedding models:
    • Pros: deterministic, fast, no model downloads, no internet required.
    • Cons: weaker semantic recall than transformer embeddings.
  • Local extractive generation vs generative LLM:
    • Pros: grounded-by-construction, citation-safe, no hallucinated synthesis.
    • Cons: less fluent and less abstractive than LLM outputs.

Security Notes

  • No secrets are committed.
  • Use .env (not tracked) for credentials.
  • mode=local performs no outbound Azure calls.
  • Review SECURITY.md for disclosure guidance.

Repository Commands

# lint + tests
.\scripts\test.ps1

# run demo end-to-end
.\scripts\demo.ps1

License

MIT (LICENSE).

About

Retrieval-Augmented Generation pipeline for Azure documentation with citation support, using Azure AI Search or Qdrant and Azure OpenAI or local extractive answering.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages