A local-first RAG system with persistent memory, exact replay and interactive learning via web interface — no retraining required.
Learn new knowledge instantly by adding and verifying content directly in the UI.
Generate once, store forever, retrieve exactly — or refine and evolve your knowledge over time.
This system separates generation from memory:
- The LLM generates content
- The system stores and curates it
- Future answers are retrieved — not hallucinated
👉 Knowledge improves over time through human feedback.
RAG Memory Bot is a local FastAPI-based prototype for a retrieval-augmented long-term memory system with focus on three core modes:
- Generate: create new content
- Replay: reproduce stored content exactly
- RAG Answer / Variation: use stored knowledge for answers or variations
The system combines Markdown storage, SQLite metadata, full-text search (FTS5), embedding-based retrieval, and a lightweight web interface.
- Persistent storage of generated content as Markdown artifacts
- Metadata management via SQLite
- Hybrid search combining:
- Full-text search (FTS5)
- Semantic search (embeddings)
- Status / feedback ranking
- Session-aware replay and variation
- Review workflow with statuses:
draft,saved,favorite,fixed,verified,archived
- Import of existing Markdown files
- Web interface for:
- Chat
- Artifact management
- Review and verification
- Data import
User → API (FastAPI) → Orchestrator
Orchestrator coordinates:
- Retrieval (FTS + vector search)
- Storage (Markdown + SQLite)
- LLM (Ollama)
(Add your screenshots here)
- Course-specific assistants
- FAQ systems
- Knowledge-based tutoring
- Internal knowledge systems
- Policy-aware assistants
- Local, privacy-friendly AI
- Story generation with memory
- Replayable narratives
- Interactive storytelling
- Persistent assistant memory
- Retrieval of past answers
- Knowledge refinement via feedback
git clone https://github.com/YOUR_USERNAME/rag-memory-bot.git
cd rag-memory-botpython3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtCheck installation:
ollama listpython3 scripts/init_db.pyuvicorn app.main:app --host 0.0.0.0 --port 8000Open in browser:
-
Generate content
→ "Tell me a story about a hedgehog in a storm" -
Verify content in the UI
→ Mark artifact asverified -
Replay
→ "Tell me the same story again"
→ Exact same output -
Ask a knowledge question
→ "What was the story about?"
→ Answer is generated via RAG using stored content
app/
api/
core/
storage/
retrieval/
llm/
indexing/
templates/
static/
data/
artifacts/
db/
docs/
Startanleitung.md
scripts/
This project is licensed under the MIT License.
Note: This project uses third-party models (e.g. via Ollama).
Their respective licenses apply separately.