Skip to content

Migrate from Gemini to OpenRouter free models + local embeddings - #1

Merged
rahulk98 merged 5 commits into
masterfrom
gemini-openrouter-migration
Aug 14, 2026
Merged

rahulk98 merged 5 commits into
masterfrom
gemini-openrouter-migration

Conversation

@rahulk98

@rahulk98 rahulk98 commented Aug 7, 2026

Copy link
Copy Markdown
Owner
  • LLM: gemini-2.5-flash -> openrouter/free (OpenRouter's free-model router),
    overridable via OPENROUTER_MODEL; auth via OPENROUTER_API_KEY
  • Embeddings: gemini-embedding-001 -> local BAAI/bge-small-en-v1.5 via
    fastembed (ONNX, in-process, zero API cost); OpenRouter has no
    embeddings API, so this removes the Google dependency entirely
  • Index must be rebuilt once (embedding dims changed 3072 -> 384)
  • Dockerfile pre-downloads the embedding model at build time
  • Add pytest suite: API endpoints, rate/daily limits, caching, SSE
    streaming, index persist/reload round-trip, plus opt-in live
    OpenRouter smoke tests (pytest -m live)

- LLM: gemini-2.5-flash -> openrouter/free (OpenRouter's free-model router),
  overridable via OPENROUTER_MODEL; auth via OPENROUTER_API_KEY
- Embeddings: gemini-embedding-001 -> local BAAI/bge-small-en-v1.5 via
  fastembed (ONNX, in-process, zero API cost); OpenRouter has no
  embeddings API, so this removes the Google dependency entirely
- Index must be rebuilt once (embedding dims changed 3072 -> 384)
- Dockerfile pre-downloads the embedding model at build time
- Add pytest suite: API endpoints, rate/daily limits, caching, SSE
  streaming, index persist/reload round-trip, plus opt-in live
  OpenRouter smoke tests (pytest -m live)
Backwards compatibility for Docker rebuilds that still ship the old
Gemini-built index (3072-dim): at startup the app compares the persisted
index's embedding dimension with the current local model's (384) and
rebuilds the index from docs/ on mismatch (or if loading fails), instead
of serving broken queries. Docker image now includes docs/ to enable
this. build_index.py no longer requires OPENROUTER_API_KEY for indexing
since embeddings are local; the key is only needed for its test query.
Live testing showed /query/stream delivered the whole answer as a single
SSE event: current llama-index's Refine synthesizer buffers the entire
LLM stream and yields it once, so as_query_engine(streaming=True) no
longer streams token-by-token. The endpoint now retrieves context via
the index retriever, builds the same QA prompt, and streams deltas from
llm.stream_complete directly (verified live: ~300 SSE events, first
token ~2.7s). Mid-stream errors now end the stream with an error event
instead of aborting silently, and partial responses are not cached.
Tightened the SSE test to require genuine multi-event streaming.
- Hybrid retrieval (retrieval.py): BM25 + vector search fused with
  reciprocal rank fusion, reranked by a local ONNX cross-encoder
  (Xenova/ms-marco-MiniLM-L-6-v2). All local, zero API cost.
- Multi-turn chat: optional session_id on /query and /query/stream;
  follow-up questions are condensed into standalone queries using the
  session history (in-memory store with TTL, size caps). Cache now keys
  on the condensed question.
- Degenerate-answer guard: free-router responses that come back as junk
  (observed: a bare moderation verdict) are retried on a re-route.
- Eval harness (evals/): golden Q&A set over committed sample docs;
  offline retrieval-quality gate (100% hit rate, threshold 90%) plus
  --live mode scoring keyword coverage, abstention, and LLM-judged
  faithfulness via OpenRouter.
- CI (GitHub Actions): pytest suite + offline eval gate + Docker build
  on every push/PR.
- Endpoints now build prompts via the shared pipeline instead of
  as_query_engine; tests extended to 28 (sessions, hybrid retrieval,
  reranker ordering, eval gate).
Shipping docs/ in the image put raw source files (which may hold more
than the indexed text, e.g. PDF metadata) into every image layer and the
container registry. The image now ships only the pre-built index, and a
build-time guard (check_index.py) fails the docker build if that index
was built with a different embedding model — surfacing the
Gemini-to-local migration at build time instead of at runtime. A missing
index is still allowed (CI builds with an empty placeholder). The
startup auto-rebuild remains as a local-dev safety net when docs/ is
present.
@rahulk98
rahulk98 merged commit bc47c26 into master Aug 14, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant