Hosted and running at rahul-krishnan.is-a.dev.
Retrieval-Augmented Generation (RAG) API that indexes documents in docs/ and answers questions using LlamaIndex. The LLM runs on OpenRouter's free-model router (openrouter/free) and everything else runs locally — so there are no per-query API charges. FastAPI provides /health, /query, and /query/stream endpoints. Containerized with Docker and deployed as an API on GCP Free tier.
Pipeline highlights
- Hybrid retrieval: BM25 keyword search fused with vector search (reciprocal rank fusion), then reranked by a local ONNX cross-encoder (
Xenova/ms-marco-MiniLM-L-6-v2) — better recall than embeddings alone, still zero API cost. - Local embeddings:
BAAI/bge-small-en-v1.5via fastembed (ONNX, in-process). - Multi-turn chat: pass a
session_idand follow-up questions are condensed into standalone queries using conversation history ("what did he do there?" → "what did Rahul do at SAP?"). - Degenerate-answer guard: the free router occasionally returns a junk response; answers are validated and retried on a re-routed model.
- Evaluated: a golden Q&A set gates retrieval quality in CI (offline) and scores answer faithfulness with an LLM judge (live) — see
evals/. - Production hardening: per-IP rate limiting, daily quota, response caching, SSE streaming, query logging.
- Python: 3.12+
- Environment:
OPENROUTER_API_KEYset (in shell or.env) — get a free key at openrouter.ai- Optional:
OPENROUTER_MODELto pin a specific model (defaultopenrouter/free, which auto-routes to an available free model). Note: do not useopenrouter/auto:free— it does not restrict routing to free models and can incur charges. - Optional:
EMBED_MODEL_NAMEto change the local embedding model (defaultBAAI/bge-small-en-v1.5). Changing it requires rebuilding the index. - Optional Sheets logging: set
SHEETS_JSON_KEY(service-account JSON)
- Optional:
Migrating from the Gemini version? The embedding model changed (3072-dim Gemini → 384-dim local BGE), so the old persisted index is incompatible. Rebuild it once with
python build_index.py(no API key needed — embeddings run locally). Safety nets: the Docker build fails fast if the shipped index doesn't match the embedding model (check_index.py), and at startup the app rebuilds a mismatched index fromdocs/when that folder is present. Raw documents are never copied into the Docker image.
# Install dependencies
pip install -r requirements.txt
# Add your documents to ./docs, then build and persist the index to ./index
python build_index.py
# Start the API server
python -m uvicorn app:app --host 0.0.0.0 --port 8000- GET
/health→{ "status": "ok", "model_loaded": boolean } - POST
/query- Request body (
top_kandsession_idoptional; send a stablesession_idto enable follow-up questions):
- Request body (
{ "query": "Tell me about Rahul in 2 sentences.", "top_k": 5, "session_id": "abc123" }- Response body:
{ "query": "...", "answer": "..." }- POST
/query/stream— same request body; responds withtext/event-streamofdata: {"token": "..."}events terminated bydata: [DONE].
# Build
docker build -t rag-resume-api .
# Run
docker run -e OPENROUTER_API_KEY=YOUR_KEY -p 8000:8000 rag-resume-apipip install pytest httpx
pytest tests/ -v # unit/API tests — LLM mocked, no key needed
python evals/run_eval.py # offline retrieval-quality eval (also runs in CI)With OPENROUTER_API_KEY set:
pytest tests/ -v -m live # live OpenRouter smoke tests
python evals/run_eval.py --live # full answer eval: keyword coverage, abstention, LLM-judged faithfulness
