Skip to content

Repository files navigation

RAG Resume API

CI

Live

Hosted and running at rahul-krishnan.is-a.dev.

Overview

Retrieval-Augmented Generation (RAG) API that indexes documents in docs/ and answers questions using LlamaIndex. The LLM runs on OpenRouter's free-model router (openrouter/free) and everything else runs locally — so there are no per-query API charges. FastAPI provides /health, /query, and /query/stream endpoints. Containerized with Docker and deployed as an API on GCP Free tier.

Pipeline highlights

  • Hybrid retrieval: BM25 keyword search fused with vector search (reciprocal rank fusion), then reranked by a local ONNX cross-encoder (Xenova/ms-marco-MiniLM-L-6-v2) — better recall than embeddings alone, still zero API cost.
  • Local embeddings: BAAI/bge-small-en-v1.5 via fastembed (ONNX, in-process).
  • Multi-turn chat: pass a session_id and follow-up questions are condensed into standalone queries using conversation history ("what did he do there?" → "what did Rahul do at SAP?").
  • Degenerate-answer guard: the free router occasionally returns a junk response; answers are validated and retried on a re-routed model.
  • Evaluated: a golden Q&A set gates retrieval quality in CI (offline) and scores answer faithfulness with an LLM judge (live) — see evals/.
  • Production hardening: per-IP rate limiting, daily quota, response caching, SSE streaming, query logging.

Prerequisites

  • Python: 3.12+
  • Environment: OPENROUTER_API_KEY set (in shell or .env) — get a free key at openrouter.ai
    • Optional: OPENROUTER_MODEL to pin a specific model (default openrouter/free, which auto-routes to an available free model). Note: do not use openrouter/auto:free — it does not restrict routing to free models and can incur charges.
    • Optional: EMBED_MODEL_NAME to change the local embedding model (default BAAI/bge-small-en-v1.5). Changing it requires rebuilding the index.
    • Optional Sheets logging: set SHEETS_JSON_KEY (service-account JSON)

Migrating from the Gemini version? The embedding model changed (3072-dim Gemini → 384-dim local BGE), so the old persisted index is incompatible. Rebuild it once with python build_index.py (no API key needed — embeddings run locally). Safety nets: the Docker build fails fast if the shipped index doesn't match the embedding model (check_index.py), and at startup the app rebuilds a mismatched index from docs/ when that folder is present. Raw documents are never copied into the Docker image.

Quickstart

# Install dependencies
pip install -r requirements.txt

# Add your documents to ./docs, then build and persist the index to ./index
python build_index.py

# Start the API server
python -m uvicorn app:app --host 0.0.0.0 --port 8000

API

  • GET /health → { "status": "ok", "model_loaded": boolean }
  • POST /query
    • Request body (top_k and session_id optional; send a stable session_id to enable follow-up questions):
{ "query": "Tell me about Rahul in 2 sentences.", "top_k": 5, "session_id": "abc123" }
  • Response body:
{ "query": "...", "answer": "..." }
  • POST /query/stream — same request body; responds with text/event-stream of data: {"token": "..."} events terminated by data: [DONE].

Docker

# Build
docker build -t rag-resume-api .

# Run
docker run -e OPENROUTER_API_KEY=YOUR_KEY -p 8000:8000 rag-resume-api

Tests & evals

pip install pytest httpx
pytest tests/ -v            # unit/API tests — LLM mocked, no key needed
python evals/run_eval.py    # offline retrieval-quality eval (also runs in CI)

With OPENROUTER_API_KEY set:

pytest tests/ -v -m live         # live OpenRouter smoke tests
python evals/run_eval.py --live  # full answer eval: keyword coverage, abstention, LLM-judged faithfulness

Screenshots

Screenshot 1 Screenshot 2

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages