Skip to content
Ankan0503Public

About

Voice-enabled Hindi RAG over a 3.43M-vector Qdrant index. Speech streams to Sarvam realtime STT as you talk, queries are answered by hybrid dense + sparse BM25 retrieval with reciprocal rank fusion, and answers stream back token-by-token through four guardrail stages. FastAPI + React. HH Goa 2026.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

🌊 RAG in GOA: Voice-Enabled Indic Retrieval-Augmented Generation

HH Goa 2026 β€” Task 2: Build a Voice-Enabled RAG Model
Dataset: ai4bharat/MSMARCO-XI (Hindi Split β€” 97,941 Queries, 953,358 Parent Passages, 3.43M Vectors)
Team: Byte Me


πŸ‘₯ Team: Byte Me

Member Socials & Links
Ankan Giri GitHub LinkedIn X Instagram
Sayan Sinha GitHub LinkedIn X Instagram

🌟 Executive Summary

RAG in GOA is a voice-enabled Retrieval-Augmented Generation (RAG) system built for Hindi. A user speaks into their browser; audio streams live to Sarvam AI's realtime STT and the transcript appears as they talk. After the user reviews and sends it, the query passes an input safety gate, is answered from a 3.43-million vector Qdrant index via Hybrid Dense + Sparse BM25 + Reciprocal Rank Fusion (RRF), and the LLM's response streams back token-by-token over the same WebSocket. A grounding check runs once the answer is complete and flags β€” rather than blocks β€” anything the retrieved context doesn't actually support, with per-stage latency measured throughout.

[Voice Input] ──► [Sarvam Realtime STT] ──► [User reviews & sends] ──► [Input Gate] ──► [Hybrid RRF Search]
                                                                                                β”‚
                                                                                                β–Ό
                                                                          [Retrieval Gate] ──► [Parent Passage DB]
                                                                                                β”‚
                                                                                                β–Ό
[Streamed to UI, token-by-token] ◄──────────────────── [Groq / Sarvam LLM]
                β”‚
                β–Ό
[Grounding Gate, checked after streaming] ──► flags the answer already on screen if unsupported

πŸ—οΈ System Architecture

flowchart TD
    subgraph Client ["Frontend (React 19 + Vite + WebAudio)"]
        A[User Voice / Microphone] -->|Raw 16kHz PCM Frames| B[WebSocket /ws/voice-rag]
        C[Text Query Input] -->|JSON Message| B
        B -->|Stream Tokens & Latency Breakdowns| D[Live UI & Source Citations]
    end

    subgraph Backend ["Backend Engine (FastAPI + Python 3.12)"]
        B --> E[Sarvam Realtime STT: saaras:v3-realtime]
        E -->|Hindi Transcript, reviewed by user before Send| F[Input Safety & Injection Guardrail]
        F -->|Clean Query| G

        subgraph HybridRetrieval ["Hybrid RRF Retrieval"]
            G[Dense Encoder: multilingual-e5-small] --> I[Qdrant Vector Server: HNSW Graph]
            H[Sparse Encoder: Qdrant/bm25] --> I
            I -->|Top-K Child Chunk IDs| J[Reciprocal Rank Fusion RRF]
            J -->|Parent IDs, deduplicated| K[(parents.sqlite - 953k rows)]
            K --> L[Context Assembler]
        end

        L --> M[Retrieval Confidence Gate: cosine floor + relative margin]
        M -->|Context Verified| N[Groq / Sarvam LLM, streamed token-by-token]
        M -->|Low Confidence / Out-of-Domain| O[Pre-compiled Safe Refusal, nothing generated]

        N -->|Tokens stream to the client as they're produced| B
        N -->|Once the full answer exists| P[Grounding Gate: lexical overlap vs context]
        P -->|Below threshold| Q[Flag already-streamed answer as unverified]
        Q --> B
        O -->|Stream Refusal| B
    end
Loading

Grounding is checked after the answer has already streamed to the browser, not before β€” the check needs the complete text, and blocking the whole stream on it would defeat the point of streaming. An answer that fails the check isn't hidden; it's dimmed and flagged in the UI with the real overlap score, so a false answer stays visible as "flagged" rather than silently disappearing.


🧩 Key Innovations & Technical Highlights

1. πŸŽ™οΈ Real-Time WebSocket Speech-to-Text

  • Captures the microphone via the WebAudio API (ScriptProcessorNode, 4096-sample buffer β€” roughly 85ms per frame at a typical 48kHz device rate) and resamples to the 16kHz mono s16le PCM Sarvam's realtime socket requires, rather than uploading a MediaRecorder WebM/Opus blob after the user stops speaking.
  • Streams those frames to Sarvam AI saaras:v3-realtime over WebSocket as the user talks, receiving partial transcripts live and a final transcript on completion. The user reviews the transcript and presses Send before anything is retrieved or generated β€” confirmed against real speech via backend/probe_realtime_stt.py, not just the docs.
  • The batch REST client (saarika:v2.5, backend/audio_stt.py) remains available as a fallback path, with retries.

2. πŸ“š Vast Multi-Granularity Adaptive Chunking Strategy

Passages in MSMARCO-XI Hindi range from 1-sentence statements to dense 2,000-character documents. A naive fixed-size chunking strategy shreds short passages and drowns key sentences in long ones.

Our pipeline implements Length-Adaptive Multi-Strategy Chunking β€” every passage is routed to whichever strategies apply to its own length, not chunked one way regardless of what it contains:

  • S1 Β· Passage: Full passage as one vector. Runs for every passage as the baseline.
  • S2 Β· Sentence: Context-prefixed, one vector per sentence, for passages of 3–6 sentences. Lets a single precise sentence surface on its own instead of being outweighed by the rest of its passage.
  • S3 Β· Sliding Window: 3-sentence windows with 1-sentence overlap, for passages of 7+ sentences, so a fact sitting at a window boundary isn't cut in half.
  • S4 Β· Semantic: Embedding-similarity boundary detection groups sentences by topic shift instead of a fixed count, also for 7+ sentence passages.
  • Parent-Child Architecture: Chunks act as children indexed in Qdrant; answers are generated from the parent passage stored in parents.sqlite, giving retrieval precision without losing surrounding context. All 10 retrieved passages per query are indexed, not only the ~7% marked "selected" in the source data β€” discarding the rest would throw away 93% of the corpus and leave over a third of queries with nothing indexed at all.
Total Indexed: 97,941 Queries | 953,358 Parent Passages | 3,433,257 Vectors in Qdrant

(Per-strategy vector counts aren't reproduced here β€” pull them from the build log's by_strategy counter or index_manifest.json if you need the exact breakdown.)

3. ⚑ Hybrid Dense + Sparse BM25 Search with RRF

  • Dense Vectors: intfloat/multilingual-e5-small (384 dimensions, scalar quantized to INT8) for deep cross-lingual semantic capture.
  • Sparse BM25: Qdrant/bm25 inverted index for exact keyword, numerical, entity, and name matches (essential for ~40% of MSMARCO entity queries).
  • Reciprocal Rank Fusion (RRF): Merges dense and sparse ranks with parent deduplication so distinct chunking tiers reinforce document relevance rather than crowding out the top-k slots.

4. πŸ›‘οΈ 4-Stage Guardrails & Hallucination Prevention

  1. Input Gate: Rejects empty/oversized queries and a small deliberately-conservative lexical filter for prompt injection and unsafe input.
  2. Retrieval Gate: An absolute cosine floor (MIN_RETRIEVAL_SCORE = 0.850, calibrated against real vs. out-of-domain query score distributions) ANDed with a relative-margin check β€” whether the top hit actually stands out from the rest of the field, not just whether it clears a fixed bar. Refuses out-of-domain questions before any LLM call runs, saving generation cost entirely.
  3. Grounding Gate: Lexical token-overlap check (MIN_GROUNDING_OVERLAP = 0.45) between the generated answer and the retrieved context β€” not a trained NLI/entailment model. Because the answer streams to the client before this check can run, a failure doesn't block the stream; it flags the already-visible answer as unverified, with the real overlap score shown in the UI.
  4. Output Gate: Checks the answer is majority-Devanagari script and free of prompt-leak markers (e.g. the system prompt's own context/instruction delimiters bleeding into the response).

πŸ“‚ Repository Structure

RagInGoa/
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ server.py              # FastAPI server (REST + WebSocket /ws/voice-rag)
β”‚   β”œβ”€β”€ retriever.py           # Hybrid dense + sparse BM25 + RRF engine
β”‚   β”œβ”€β”€ build_index_gpu.py     # GPU-accelerated indexing pipeline
β”‚   β”œβ”€β”€ stt_realtime.py        # Sarvam Realtime WebSocket streaming client
β”‚   β”œβ”€β”€ audio_stt.py           # Sarvam REST STT client with exponential backoff
β”‚   β”œβ”€β”€ guardrails.py          # 4-stage input, retrieval, grounding, and output guardrails
β”‚   β”œβ”€β”€ llm.py                 # Dual LLM provider abstraction (Groq LPUs + Sarvam Indic)
β”‚   β”œβ”€β”€ profiling.py           # Microsecond latency tracker
β”‚   β”œβ”€β”€ benchmark.py           # Evaluation & percentile benchmarking suite
β”‚   β”œβ”€β”€ index_manifest.json    # NOT in git (gitignored). Index metadata, config fingerprint β€”
β”‚   β”‚                          #   generated by build_index_gpu.py, placed at deploy time.
β”‚   β”œβ”€β”€ parents.sqlite         # NOT in git (gitignored, *.sqlite). Parent passage storage
β”‚   β”‚                          #   (953k rows, ~1.4GB) β€” also placed at deploy time.
β”‚   └── requirements.txt       # Backend Python dependencies
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ components/
β”‚   β”‚   β”‚   β”œβ”€β”€ AskHero/       # Center microphone & real-time waveform interface
β”‚   β”‚   β”‚   β”œβ”€β”€ AnswerPanel/   # Streamed answer, latency gauges, source citations
β”‚   β”‚   β”‚   β”œβ”€β”€ AboutSection/  # Team, project description, chunking strategy writeup
β”‚   β”‚   β”‚   β”œβ”€β”€ LeftSidebar/   # Navigation (Ask / Insights / About)
β”‚   β”‚   β”‚   └── BottomAskSection/ # Capability highlights
β”‚   β”‚   β”œβ”€β”€ context/
β”‚   β”‚   β”‚   └── RagContext.tsx # WebAudio stream & WebSocket state manager
β”‚   β”‚   └── App.tsx            # Main responsive layout container
β”‚   β”œβ”€β”€ package.json
β”‚   └── vite.config.ts
β”œβ”€β”€ Dockerfile                 # Multi-stage container build
β”œβ”€β”€ docker-compose.yml         # Production stack (Qdrant + FastAPI + Frontend)
β”œβ”€β”€ entrypoint.sh              # Container initialization script
└── README.md

πŸš€ Quick Start Guide

Prerequisites

  • Python 3.12 (matches the Docker image; other 3.x versions are untested)
  • Node.js 20+ & npm
  • API Keys:
    • SARVAM_API_KEY (from Sarvam AI) β€” also used for realtime STT
    • GROQ_API_KEY (from Groq Console)
    • A pre-built index (parents.sqlite + index_manifest.json) placed in backend/, and either a local qdrant_data/ or a running Qdrant server β€” the app will not start without one. See backend/build_index_gpu.py.

1. Backend Setup

# .env.example lives at the repo root, not inside backend/. Two separate .env
# files exist by design: docker-compose reads one at the repo root; running
# server.py directly (this section) reads backend/.env instead.
cd backend
cp ../.env.example .env
# Edit backend/.env with your API keys (SARVAM_API_KEY, GROQ_API_KEY,
# QDRANT_API_KEY if pointing at a Qdrant server instead of local file mode)

# Create & activate virtual environment
python -m venv venv
# On Windows:
.\venv\Scripts\Activate.ps1
# On Linux/macOS:
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Start FastAPI dev server
python -m uvicorn server:app --host 0.0.0.0 --port 8000 --reload

2. Frontend Setup

# Navigate to frontend (in a new terminal)
cd frontend

# Install dependencies
npm install

# Start Vite dev server
npm run dev

Open your browser at http://localhost:5173.


3. Docker Compose (Production Deployment)

To deploy the entire production stack (FastAPI Backend + Built Frontend + Qdrant Server):

# Start all services in the background
docker-compose up -d --build

# View logs
docker-compose logs -f

πŸ“Š Benchmarks & Latency Profiling

Retrieval: local file mode vs. Qdrant server

Qdrant's local file mode does brute-force NumPy search and ignores quantization entirely; the migration to a real Qdrant server (HNSW + int8 quantization) was made specifically because of this measured gap. Measured on the 680K-vector index (20K-query subset, before the full 97,941-query rebuild):

Metric Local file mode Qdrant server Speedup
Retrieval P50 11,411 ms 75.7 ms 151Γ—
Retrieval P100 26,744 ms 161 ms 166Γ—
Startup time ~75 min < 40 s β€”

Full pipeline, by model β€” full 97,941-query index (3.43M vectors)

Measured via backend/benchmark.py --compare-models against the live pipeline (retrieval + both guardrails + generation) on the deployed, full-scale index β€” not the 680K-vector subset above. Two storage fixes were applied to the live Qdrant collection before this run: the HNSW graph and the BM25 sparse index were both switched from Qdrant's default memory-mapped (disk) storage to RAM-resident. At 680K vectors this distinction didn't matter β€” the graph fit in the OS page cache regardless. At 3.43M vectors it did: retrieval P50 sat at 120–130ms with the graph on disk, every model, regardless of which one answered β€” confirming the cost was architectural, not model-dependent, before the fix below.

Model Retrieval P50/P100 Generation P50/P100 End-to-end P50/P100 Refused Recall
qwen/qwen3-32b 132.6 / 4,235 msΒΉ 562.5 / 1,220 ms 689.7 / 5,650 msΒΉ 2/30 (retrieval gate only) 0.280
openai/gpt-oss-20b 94.0 / 166.0 ms 569.0 / 825.7 ms 667.0 / 910.1 ms 10/30 (9 grounding, 1 retrieval) 0.276
groq/compound-mini 90.8 / 170.3 ms 574.3 / 1,615 ms 665.3 / 1,712 ms 2/30 (retrieval gate only) 0.348

ΒΉ qwen's P100 here is a single one-time outlier on the first query of this run β€” the first search to touch a just-migrated HNSW segment paid a real page-fault cost once. Every query after it in this same run, and every query in the two model runs that followed, ran clean; qwen's P70/P90 for this run were 169ms/195ms, in line with the other two models. Not reproducible on a re-run against a settled index.

gpt-oss-20b's grounding-gate refusal rate (9/30, ~30%) is real and reproduces across multiple independent runs on this data β€” the model generates an answer, then the post-generation lexical-overlap check rejects it. Whether this is a genuine grounding failure or a format mismatch with the overlap check has not been investigated (see Guardrails above). compound-mini and qwen/qwen3-32b show 0% grounding-gate refusals across every run measured this session; both refusals shown for each are the retrieval gate declining to answer before generation runs at all.


πŸ“œ License

This project is licensed under the Apache License 2.0. Developed for HH Goa 2026.

About

Voice-enabled Hindi RAG over a 3.43M-vector Qdrant index. Speech streams to Sarvam realtime STT as you talk, queries are answered by hybrid dense + sparse BM25 retrieval with reciprocal rank fusion, and answers stream back token-by-token through four guardrail stages. FastAPI + React. HH Goa 2026.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages