A RAG-based document intelligence system that lets you have a real conversation with any document you upload.
Upload a PDF, ask questions in natural language, and get precise, source-grounded answers — no hallucinations, just what's actually in the document. Supports multi-turn conversation so you can ask follow-up questions naturally.
- Frontend — Streamlit (dark/light theme)
- Embeddings — HuggingFace
all-MiniLM-L6-v2(384-dim) - Vector Store — ChromaDB (local persistent)
- Retrieval — Maximal Marginal Relevance (MMR)
- LLM Backends — Groq (LLaMA 3.1-8B), Gemini 2.5 Flash, Mistral 7B
- Framework — LangChain
- Multi-turn conversational QA with conversation memory
- MMR retrieval for diverse, non-redundant context
- Confidence scoring with visual indicator
- Readability analysis (Flesch Reading Ease)
- Auto-generated document summary
- Answer pinning and session export
- Persistent vector index (index once, reuse across sessions)
- Python 3.10+
- API keys for your preferred LLM backend (Groq / Gemini / Mistral)
git clone https://github.com/tahirshamim/folio
cd folio
pip install -r requirements.txtCreate a .env file in the root directory:
GROQ_API_KEY=your_groq_key
GOOGLE_API_KEY=your_gemini_key
MISTRAL_API_KEY=your_mistral_keystreamlit run app.py- Ingest — Document is parsed and split into 1000-char chunks with 200-char overlap
- Embed — Each chunk is embedded using
all-MiniLM-L6-v2and stored in ChromaDB - Retrieve — Query is embedded and MMR search fetches the 5 most relevant, diverse passages
- Generate — Retrieved passages + conversation history are sent to the selected LLM with a strict grounding prompt
folio/
├── app.py # Main Streamlit application
├── requirements.txt
├── .env.example
└── chroma_store/ # Persistent vector index (auto-generated)
- Multi-document RAG
- Cross-encoder re-ranking
- REST API mode (FastAPI)
- RAGAS evaluation dashboard
- User authentication
- Tahir Bin Shamim
MIT