HH Goa 2026 β Task 2: Build a Voice-Enabled RAG Model
Dataset:ai4bharat/MSMARCO-XI(Hindi Split β 97,941 Queries, 953,358 Parent Passages, 3.43M Vectors)
Team: Byte Me
| Member | Socials & Links |
|---|---|
| Ankan Giri | |
| Sayan Sinha |
RAG in GOA is a voice-enabled Retrieval-Augmented Generation (RAG) system built for Hindi. A user speaks into their browser; audio streams live to Sarvam AI's realtime STT and the transcript appears as they talk. After the user reviews and sends it, the query passes an input safety gate, is answered from a 3.43-million vector Qdrant index via Hybrid Dense + Sparse BM25 + Reciprocal Rank Fusion (RRF), and the LLM's response streams back token-by-token over the same WebSocket. A grounding check runs once the answer is complete and flags β rather than blocks β anything the retrieved context doesn't actually support, with per-stage latency measured throughout.
[Voice Input] βββΊ [Sarvam Realtime STT] βββΊ [User reviews & sends] βββΊ [Input Gate] βββΊ [Hybrid RRF Search]
β
βΌ
[Retrieval Gate] βββΊ [Parent Passage DB]
β
βΌ
[Streamed to UI, token-by-token] βββββββββββββββββββββ [Groq / Sarvam LLM]
β
βΌ
[Grounding Gate, checked after streaming] βββΊ flags the answer already on screen if unsupported
flowchart TD
subgraph Client ["Frontend (React 19 + Vite + WebAudio)"]
A[User Voice / Microphone] -->|Raw 16kHz PCM Frames| B[WebSocket /ws/voice-rag]
C[Text Query Input] -->|JSON Message| B
B -->|Stream Tokens & Latency Breakdowns| D[Live UI & Source Citations]
end
subgraph Backend ["Backend Engine (FastAPI + Python 3.12)"]
B --> E[Sarvam Realtime STT: saaras:v3-realtime]
E -->|Hindi Transcript, reviewed by user before Send| F[Input Safety & Injection Guardrail]
F -->|Clean Query| G
subgraph HybridRetrieval ["Hybrid RRF Retrieval"]
G[Dense Encoder: multilingual-e5-small] --> I[Qdrant Vector Server: HNSW Graph]
H[Sparse Encoder: Qdrant/bm25] --> I
I -->|Top-K Child Chunk IDs| J[Reciprocal Rank Fusion RRF]
J -->|Parent IDs, deduplicated| K[(parents.sqlite - 953k rows)]
K --> L[Context Assembler]
end
L --> M[Retrieval Confidence Gate: cosine floor + relative margin]
M -->|Context Verified| N[Groq / Sarvam LLM, streamed token-by-token]
M -->|Low Confidence / Out-of-Domain| O[Pre-compiled Safe Refusal, nothing generated]
N -->|Tokens stream to the client as they're produced| B
N -->|Once the full answer exists| P[Grounding Gate: lexical overlap vs context]
P -->|Below threshold| Q[Flag already-streamed answer as unverified]
Q --> B
O -->|Stream Refusal| B
end
Grounding is checked after the answer has already streamed to the browser, not before β the check needs the complete text, and blocking the whole stream on it would defeat the point of streaming. An answer that fails the check isn't hidden; it's dimmed and flagged in the UI with the real overlap score, so a false answer stays visible as "flagged" rather than silently disappearing.
- Captures the microphone via the WebAudio API (
ScriptProcessorNode, 4096-sample buffer β roughly 85ms per frame at a typical 48kHz device rate) and resamples to the 16kHz monos16lePCM Sarvam's realtime socket requires, rather than uploading a MediaRecorder WebM/Opus blob after the user stops speaking. - Streams those frames to Sarvam AI
saaras:v3-realtimeover WebSocket as the user talks, receiving partial transcripts live and a final transcript on completion. The user reviews the transcript and presses Send before anything is retrieved or generated β confirmed against real speech viabackend/probe_realtime_stt.py, not just the docs. - The batch REST client (
saarika:v2.5,backend/audio_stt.py) remains available as a fallback path, with retries.
Passages in MSMARCO-XI Hindi range from 1-sentence statements to dense 2,000-character documents. A naive fixed-size chunking strategy shreds short passages and drowns key sentences in long ones.
Our pipeline implements Length-Adaptive Multi-Strategy Chunking β every passage is routed to whichever strategies apply to its own length, not chunked one way regardless of what it contains:
- S1 Β· Passage: Full passage as one vector. Runs for every passage as the baseline.
- S2 Β· Sentence: Context-prefixed, one vector per sentence, for passages of 3β6 sentences. Lets a single precise sentence surface on its own instead of being outweighed by the rest of its passage.
- S3 Β· Sliding Window: 3-sentence windows with 1-sentence overlap, for passages of 7+ sentences, so a fact sitting at a window boundary isn't cut in half.
- S4 Β· Semantic: Embedding-similarity boundary detection groups sentences by topic shift instead of a fixed count, also for 7+ sentence passages.
- Parent-Child Architecture: Chunks act as children indexed in Qdrant; answers are generated from the parent passage stored in
parents.sqlite, giving retrieval precision without losing surrounding context. All 10 retrieved passages per query are indexed, not only the ~7% marked "selected" in the source data β discarding the rest would throw away 93% of the corpus and leave over a third of queries with nothing indexed at all.
Total Indexed: 97,941 Queries | 953,358 Parent Passages | 3,433,257 Vectors in Qdrant
(Per-strategy vector counts aren't reproduced here β pull them from the build log's by_strategy counter or index_manifest.json if you need the exact breakdown.)
- Dense Vectors:
intfloat/multilingual-e5-small(384 dimensions, scalar quantized to INT8) for deep cross-lingual semantic capture. - Sparse BM25:
Qdrant/bm25inverted index for exact keyword, numerical, entity, and name matches (essential for ~40% of MSMARCO entity queries). - Reciprocal Rank Fusion (RRF): Merges dense and sparse ranks with parent deduplication so distinct chunking tiers reinforce document relevance rather than crowding out the top-k slots.
- Input Gate: Rejects empty/oversized queries and a small deliberately-conservative lexical filter for prompt injection and unsafe input.
- Retrieval Gate: An absolute cosine floor (
MIN_RETRIEVAL_SCORE = 0.850, calibrated against real vs. out-of-domain query score distributions) ANDed with a relative-margin check β whether the top hit actually stands out from the rest of the field, not just whether it clears a fixed bar. Refuses out-of-domain questions before any LLM call runs, saving generation cost entirely. - Grounding Gate: Lexical token-overlap check (
MIN_GROUNDING_OVERLAP = 0.45) between the generated answer and the retrieved context β not a trained NLI/entailment model. Because the answer streams to the client before this check can run, a failure doesn't block the stream; it flags the already-visible answer as unverified, with the real overlap score shown in the UI. - Output Gate: Checks the answer is majority-Devanagari script and free of prompt-leak markers (e.g. the system prompt's own context/instruction delimiters bleeding into the response).
RagInGoa/
βββ backend/
β βββ server.py # FastAPI server (REST + WebSocket /ws/voice-rag)
β βββ retriever.py # Hybrid dense + sparse BM25 + RRF engine
β βββ build_index_gpu.py # GPU-accelerated indexing pipeline
β βββ stt_realtime.py # Sarvam Realtime WebSocket streaming client
β βββ audio_stt.py # Sarvam REST STT client with exponential backoff
β βββ guardrails.py # 4-stage input, retrieval, grounding, and output guardrails
β βββ llm.py # Dual LLM provider abstraction (Groq LPUs + Sarvam Indic)
β βββ profiling.py # Microsecond latency tracker
β βββ benchmark.py # Evaluation & percentile benchmarking suite
β βββ index_manifest.json # NOT in git (gitignored). Index metadata, config fingerprint β
β β # generated by build_index_gpu.py, placed at deploy time.
β βββ parents.sqlite # NOT in git (gitignored, *.sqlite). Parent passage storage
β β # (953k rows, ~1.4GB) β also placed at deploy time.
β βββ requirements.txt # Backend Python dependencies
βββ frontend/
β βββ src/
β β βββ components/
β β β βββ AskHero/ # Center microphone & real-time waveform interface
β β β βββ AnswerPanel/ # Streamed answer, latency gauges, source citations
β β β βββ AboutSection/ # Team, project description, chunking strategy writeup
β β β βββ LeftSidebar/ # Navigation (Ask / Insights / About)
β β β βββ BottomAskSection/ # Capability highlights
β β βββ context/
β β β βββ RagContext.tsx # WebAudio stream & WebSocket state manager
β β βββ App.tsx # Main responsive layout container
β βββ package.json
β βββ vite.config.ts
βββ Dockerfile # Multi-stage container build
βββ docker-compose.yml # Production stack (Qdrant + FastAPI + Frontend)
βββ entrypoint.sh # Container initialization script
βββ README.md
- Python 3.12 (matches the Docker image; other 3.x versions are untested)
- Node.js 20+ & npm
- API Keys:
SARVAM_API_KEY(from Sarvam AI) β also used for realtime STTGROQ_API_KEY(from Groq Console)- A pre-built index (
parents.sqlite+index_manifest.json) placed inbackend/, and either a localqdrant_data/or a running Qdrant server β the app will not start without one. Seebackend/build_index_gpu.py.
# .env.example lives at the repo root, not inside backend/. Two separate .env
# files exist by design: docker-compose reads one at the repo root; running
# server.py directly (this section) reads backend/.env instead.
cd backend
cp ../.env.example .env
# Edit backend/.env with your API keys (SARVAM_API_KEY, GROQ_API_KEY,
# QDRANT_API_KEY if pointing at a Qdrant server instead of local file mode)
# Create & activate virtual environment
python -m venv venv
# On Windows:
.\venv\Scripts\Activate.ps1
# On Linux/macOS:
source venv/bin/activate
# Install dependencies
pip install -r requirements.txt
# Start FastAPI dev server
python -m uvicorn server:app --host 0.0.0.0 --port 8000 --reload# Navigate to frontend (in a new terminal)
cd frontend
# Install dependencies
npm install
# Start Vite dev server
npm run devOpen your browser at http://localhost:5173.
To deploy the entire production stack (FastAPI Backend + Built Frontend + Qdrant Server):
# Start all services in the background
docker-compose up -d --build
# View logs
docker-compose logs -fQdrant's local file mode does brute-force NumPy search and ignores quantization entirely; the migration to a real Qdrant server (HNSW + int8 quantization) was made specifically because of this measured gap. Measured on the 680K-vector index (20K-query subset, before the full 97,941-query rebuild):
| Metric | Local file mode | Qdrant server | Speedup |
|---|---|---|---|
| Retrieval P50 | 11,411 ms | 75.7 ms | 151Γ |
| Retrieval P100 | 26,744 ms | 161 ms | 166Γ |
| Startup time | ~75 min | < 40 s | β |
Measured via backend/benchmark.py --compare-models against the live pipeline
(retrieval + both guardrails + generation) on the deployed, full-scale index β
not the 680K-vector subset above. Two storage fixes were applied to the live
Qdrant collection before this run: the HNSW graph and the BM25 sparse index
were both switched from Qdrant's default memory-mapped (disk) storage to
RAM-resident. At 680K vectors this distinction didn't matter β the graph fit
in the OS page cache regardless. At 3.43M vectors it did: retrieval P50 sat
at 120β130ms with the graph on disk, every model, regardless of which one
answered β confirming the cost was architectural, not model-dependent, before
the fix below.
| Model | Retrieval P50/P100 | Generation P50/P100 | End-to-end P50/P100 | Refused | Recall |
|---|---|---|---|---|---|
qwen/qwen3-32b |
132.6 / 4,235 msΒΉ | 562.5 / 1,220 ms | 689.7 / 5,650 msΒΉ | 2/30 (retrieval gate only) | 0.280 |
openai/gpt-oss-20b |
94.0 / 166.0 ms | 569.0 / 825.7 ms | 667.0 / 910.1 ms | 10/30 (9 grounding, 1 retrieval) | 0.276 |
groq/compound-mini |
90.8 / 170.3 ms | 574.3 / 1,615 ms | 665.3 / 1,712 ms | 2/30 (retrieval gate only) | 0.348 |
ΒΉ qwen's P100 here is a single one-time outlier on the first query of this run β the first search to touch a just-migrated HNSW segment paid a real page-fault cost once. Every query after it in this same run, and every query in the two model runs that followed, ran clean; qwen's P70/P90 for this run were 169ms/195ms, in line with the other two models. Not reproducible on a re-run against a settled index.
gpt-oss-20b's grounding-gate refusal rate (9/30, ~30%) is real and
reproduces across multiple independent runs on this data β the model
generates an answer, then the post-generation lexical-overlap check rejects
it. Whether this is a genuine grounding failure or a format mismatch with
the overlap check has not been investigated (see Guardrails above).
compound-mini and qwen/qwen3-32b show 0% grounding-gate refusals across
every run measured this session; both refusals shown for each are the
retrieval gate declining to answer before generation runs at all.
This project is licensed under the Apache License 2.0. Developed for HH Goa 2026.