Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Folio — Document Intelligence System

A RAG-based document intelligence system that lets you have a real conversation with any document you upload.

What it does

Upload a PDF, ask questions in natural language, and get precise, source-grounded answers — no hallucinations, just what's actually in the document. Supports multi-turn conversation so you can ask follow-up questions naturally.

Tech Stack

  • Frontend — Streamlit (dark/light theme)
  • Embeddings — HuggingFace all-MiniLM-L6-v2 (384-dim)
  • Vector Store — ChromaDB (local persistent)
  • Retrieval — Maximal Marginal Relevance (MMR)
  • LLM Backends — Groq (LLaMA 3.1-8B), Gemini 2.5 Flash, Mistral 7B
  • Framework — LangChain

Features

  • Multi-turn conversational QA with conversation memory
  • MMR retrieval for diverse, non-redundant context
  • Confidence scoring with visual indicator
  • Readability analysis (Flesch Reading Ease)
  • Auto-generated document summary
  • Answer pinning and session export
  • Persistent vector index (index once, reuse across sessions)

Getting Started

Prerequisites

  • Python 3.10+
  • API keys for your preferred LLM backend (Groq / Gemini / Mistral)

Installation

git clone https://github.com/tahirshamim/folio
cd folio
pip install -r requirements.txt

Setup

Create a .env file in the root directory:

GROQ_API_KEY=your_groq_key
GOOGLE_API_KEY=your_gemini_key
MISTRAL_API_KEY=your_mistral_key

Run

streamlit run app.py

How it works

  1. Ingest — Document is parsed and split into 1000-char chunks with 200-char overlap
  2. Embed — Each chunk is embedded using all-MiniLM-L6-v2 and stored in ChromaDB
  3. Retrieve — Query is embedded and MMR search fetches the 5 most relevant, diverse passages
  4. Generate — Retrieved passages + conversation history are sent to the selected LLM with a strict grounding prompt

Project Structure

folio/
├── app.py              # Main Streamlit application
├── requirements.txt
├── .env.example
└── chroma_store/       # Persistent vector index (auto-generated)

Roadmap

  • Multi-document RAG
  • Cross-encoder re-ranking
  • REST API mode (FastAPI)
  • RAGAS evaluation dashboard
  • User authentication

Team

  • Tahir Bin Shamim

License

MIT

About

Rag based Document Intelligence system

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages