Skip to content

Repository files navigation

🦙 Loco-RAG · Local LLM Chat with RAG

A privacy-first, fully-local chat app for self-hosted LLMs — with document Q&A, agents, memory, web search, and voice.

Talk to Ollama & LM Studio models, chat with your own documents, and let an agent search them for you — all on your machine. Nothing ever leaves the box.

CI License: MIT Python React FastAPI Docker Privacy

Why? Cloud chat apps send your conversations and documents to someone else's servers. Loco-RAG gives you the same ChatGPT-class experience — streaming, RAG, citations, agents, voice — running entirely against a local model. No API keys, no telemetry, no data leaving home.

Animated flow: query → embed → vector store → RRF fusion → LLM → cited answer

✨ Highlights

💬 Streaming chat SSE, markdown + code, stop / edit-and-rerun, tokens/sec
📚 Deep RAG Hybrid retrieval (dense + BM25 → RRF), optional rerank, MMR diversification, inline citations
🗂️ Vector store ChromaDB or LanceDB — swap with one env var
📎 Talk to this doc Attach a file to a single message for ephemeral RAG
🤖 Agent mode Model decides when to search your docs / the web
🧠 Memory Older turns summarized to stay in the context window
🌐 Web search DuckDuckGo or self-hosted SearXNG, cited
🎙️ Voice I/O Local speech-to-text (faster-whisper) + read-aloud

Features

  • Streaming chat over SSE with stop / edit-and-rerun, markdown + syntax-highlighted code.
  • Multi-conversation sidebar (create, rename via auto-title, pin, search, delete).
  • Model switcher that auto-discovers installed Ollama / LM Studio models (vision models flagged).
  • RAG / document Q&A — drag-and-drop PDF, DOCX, TXT, MD, CSV, code. Token-aware chunking, hybrid retrieval (dense + BM25 fused with RRF), optional cross-encoder reranking, optional MMR diversification (ENABLE_MMR) so the context isn't several near-duplicate passages, and inline citations with an expandable source panel.
  • Collections / knowledge bases — group docs and scope a chat to a collection.
  • "Talk to this doc" — attach a file to a single message (paperclip) for ephemeral RAG over just that file, without adding it to a persistent collection.
  • Agentic tool-calling — toggle Agent and the model decides for itself when to call the document-search and web-search tools, chaining multiple rounds before answering.
  • Voice in / out — dictate with local speech-to-text (faster-whisper) and have replies read aloud with the browser's built-in speech synthesis.
  • Conversation memory — older turns are summarized to stay within the context window.
  • Web search — per-message toggle (DuckDuckGo out of the box, or self-hosted SearXNG).
  • Pluggable vector store — ChromaDB (default) or LanceDB, switch with one env var.
  • Local embeddings — fastembed (bundled) or an Ollama embedding model. No data leaves the box.
  • Themes, per-conversation system prompt and sampling params, tokens/sec readout.

Architecture

Loco-RAG architecture: React frontend talks SSE to a FastAPI backend, which talks to Ollama/LM Studio, SQLite, Chroma/LanceDB, and DuckDuckGo/SearXNG

React (Vite) ──SSE──> FastAPI ──> Ollama / LM Studio   (chat, OpenAI /v1 API)
                          ├──────> SQLite               (conversations, messages, docs)
                          ├──────> Chroma / LanceDB     (vectors)
                          └──────> DuckDuckGo / SearXNG (web search)

Key seams: services/llm (provider abstraction), services/vectorstore (Chroma/LanceDB factory), services/rag.py (hybrid retrieval + the one context-builder shared by RAG, web, and memory).

RAG pipeline: documents → parse → chunk → embed → vector store (ingest); user query → dense search + BM25 → RRF fusion → rerank → context → LLM answer with citations (retrieve)

Quick start (Docker)

Prereqs: Docker, and a local LLM runtime. For Ollama:

ollama pull llama3.1
ollama pull nomic-embed-text   # only if EMBED_PROVIDER=ollama

Then:

cp .env.example .env
docker compose up --build

Seed the default collection with the bundled sample document so RAG works immediately (ask "How tall is the Eiffel Tower?" with Documents enabled):

docker compose exec backend python -m scripts.seed      # or, in local dev: python -m scripts.seed

Enable the private SearXNG search engine instead of DuckDuckGo:

docker compose --profile websearch up --build   # then set WEBSEARCH_PROVIDER=searxng

Local dev (no Docker)

Backend

cd backend
python -m venv .venv
.venv\Scripts\activate        # Windows
source .venv/bin/activate     # macOS / Linux
pip install -e ".[dev]"
uvicorn app.main:app --reload --port 8000

Frontend

cd frontend
npm install
npm run dev      # http://localhost:5173 (proxies /api to :8000)

Switching the vector store

Set VECTOR_BACKEND=chroma or VECTOR_BACKEND=lancedb in .env, restart, and re-ingest your documents. Both implement the same VectorStore interface in backend/app/services/vectorstore/.

Tests

cd backend
pytest          # unit tests for chunking, RRF fusion, MMR, retrieval, memory,
                # agent, web search, and the retrieval-eval metrics/harness

Retrieval evaluation

A retrieval-quality harness scores whether the pipeline actually surfaces the right document. It ships a small labelled corpus (backend/eval/) and reports hit-rate, MRR, recall, precision, MAP and nDCG.

cd backend
python -m scripts.eval_retrieval             # ingest eval corpus, score retrieval
python -m scripts.eval_retrieval --k 3 --json report.json

The metrics (app/eval/metrics.py) and the aggregation harness (app/eval/harness.py) are pure and covered by pytest; the runner exercises the live embedder + vector store. Point it at your own --corpus/--dataset to guard your knowledge base against retrieval regressions.

Configuration

All settings live in .env (see .env.example) and map to backend/app/config.py.

Roadmap / ideas (deferred)

OCR for scanned PDFs · multi-modal (image) RAG · a UI dashboard over the retrieval-eval metrics · streaming tool-call traces in the UI · single-user PIN auth.

About

Privacy-first local LLM chat with RAG, agents, memory, web search & voice. Ollama/LM Studio + FastAPI + React. Hybrid retrieval, citations, ChromaDB/LanceDB. 100% local.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages