Repository navigation
Migrate from Gemini to OpenRouter free models + local embeddings - #1
Merged
Merged
Conversation
- LLM: gemini-2.5-flash -> openrouter/free (OpenRouter's free-model router), overridable via OPENROUTER_MODEL; auth via OPENROUTER_API_KEY - Embeddings: gemini-embedding-001 -> local BAAI/bge-small-en-v1.5 via fastembed (ONNX, in-process, zero API cost); OpenRouter has no embeddings API, so this removes the Google dependency entirely - Index must be rebuilt once (embedding dims changed 3072 -> 384) - Dockerfile pre-downloads the embedding model at build time - Add pytest suite: API endpoints, rate/daily limits, caching, SSE streaming, index persist/reload round-trip, plus opt-in live OpenRouter smoke tests (pytest -m live)
Backwards compatibility for Docker rebuilds that still ship the old Gemini-built index (3072-dim): at startup the app compares the persisted index's embedding dimension with the current local model's (384) and rebuilds the index from docs/ on mismatch (or if loading fails), instead of serving broken queries. Docker image now includes docs/ to enable this. build_index.py no longer requires OPENROUTER_API_KEY for indexing since embeddings are local; the key is only needed for its test query.
Live testing showed /query/stream delivered the whole answer as a single SSE event: current llama-index's Refine synthesizer buffers the entire LLM stream and yields it once, so as_query_engine(streaming=True) no longer streams token-by-token. The endpoint now retrieves context via the index retriever, builds the same QA prompt, and streams deltas from llm.stream_complete directly (verified live: ~300 SSE events, first token ~2.7s). Mid-stream errors now end the stream with an error event instead of aborting silently, and partial responses are not cached. Tightened the SSE test to require genuine multi-event streaming.
- Hybrid retrieval (retrieval.py): BM25 + vector search fused with reciprocal rank fusion, reranked by a local ONNX cross-encoder (Xenova/ms-marco-MiniLM-L-6-v2). All local, zero API cost. - Multi-turn chat: optional session_id on /query and /query/stream; follow-up questions are condensed into standalone queries using the session history (in-memory store with TTL, size caps). Cache now keys on the condensed question. - Degenerate-answer guard: free-router responses that come back as junk (observed: a bare moderation verdict) are retried on a re-route. - Eval harness (evals/): golden Q&A set over committed sample docs; offline retrieval-quality gate (100% hit rate, threshold 90%) plus --live mode scoring keyword coverage, abstention, and LLM-judged faithfulness via OpenRouter. - CI (GitHub Actions): pytest suite + offline eval gate + Docker build on every push/PR. - Endpoints now build prompts via the shared pipeline instead of as_query_engine; tests extended to 28 (sessions, hybrid retrieval, reranker ordering, eval gate).
Shipping docs/ in the image put raw source files (which may hold more than the indexed text, e.g. PDF metadata) into every image layer and the container registry. The image now ships only the pre-built index, and a build-time guard (check_index.py) fails the docker build if that index was built with a different embedding model — surfacing the Gemini-to-local migration at build time instead of at runtime. A missing index is still allowed (CI builds with an empty placeholder). The startup auto-rebuild remains as a local-dev safety net when docs/ is present.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
overridable via OPENROUTER_MODEL; auth via OPENROUTER_API_KEY
fastembed (ONNX, in-process, zero API cost); OpenRouter has no
embeddings API, so this removes the Google dependency entirely
streaming, index persist/reload round-trip, plus opt-in live
OpenRouter smoke tests (pytest -m live)