Skip to content

Repository files navigation

πŸš€ Enterprise Financial RAG Pipeline & Foundations

LLM Models Backend VectorDB Eval Deployment

A production-style, two-stage Retrieval-Augmented Generation (RAG) architecture built to parse, retrieve, and synthesize financial insights from SEC 10-K reports using the virattt/financial-qa-10K dataset.

This repository demonstrates the transition from theoretical AI research to highly scalable backend ML systems. It pairs a mathematical evaluation sandbox with an asynchronous, containerized REST API, bypassing bloated frameworks in favor of custom, deterministic, and highly optimized Python architectures powered by Gemini Flash and local BAAI embedding models.

πŸ“‘ Table of Contents


πŸ“Š Tech Stack

  • AI/ML & LLMs: Google Gemini 3.5 Flash / Flash-Lite, SentenceTransformers (BAAI/bge-large-en-v1.5), Cross-Encoders (BAAI/bge-reranker-base)
  • Evaluation: RAGAS (Context Precision, Answer Relevancy, Faithfulness), Scikit-Learn (t-SNE)
  • Backend: Python 3.12, FastAPI, Uvicorn, Pydantic, AsyncIO, Tenacity
  • Data & Storage: ChromaDB (Cosine HNSW Index), HuggingFace Datasets (virattt/financial-qa-10K)
  • DevOps & UI: Docker, Docker Compose, Streamlit, Plotly

πŸ—οΈ The Dual-Core Architecture

Building LLM applications requires separating the evaluation environment from the production runtime. This repository models that exact enterprise lifecycle.

Core 1: The Laboratory (01_rag_foundations)

Before serving users, an AI system must be mathematically validated. Core 1 is the sandbox where we prove the pipeline to benchmark retrieval strategies, determine optimal chunk sizes, and quantify model hallucinations before writing production API code.

  • πŸ“Š Offline IR Metrics: Zero-framework implementations of classic Information Retrieval benchmarksβ€”calculating Mean Precision@3, Mean Hit Rate@3 (Recall proxy), Mean Reciprocal Rank (MRR), and nDCG@5 against ground-truth keywords.
  • πŸ€– Custom LLM-as-a-Judge: Lightweight, zero-framework evaluation engine using Gemini structured outputs (RAGJudgeVerdict) to score Faithfulness (1–5), Answer Relevancy (1–5), and generate qualitative critiques alongside lexical Keyword Coverage tracking.
  • πŸ›‘οΈ Resilient API Orchestration: Rate-limit protection with exponential backoff retries (tenacity), request delay pacing, and incremental disk caching (eval_results_cache.json) to prevent data loss.
  • πŸ—ΊοΈ Dimensionality Reduction: 2D and 3D t-SNE projections using Scikit-Learn and interactive Plotly charts to visualize semantic embedding clusters across document categories.
graph TD
    A[Raw Documents] -->|Gemini JSON| B(Semantic Chunking)
    B --> C[SentenceTransformer\nBAAI/bge-large]
    C -->|Embeddings| D[(ChromaDB)]
    
    Q[User Query] --> E[Vector Search\nRecall]
    E -->|Top-K| D
    D -->|Docs| F[Cross-Encoder\nReranking]
    F -->|Ranked Docs| G[Gemini LLM]
    G --> H[Answer]
    
    H --> I{Evaluation Pipeline}
    I -->|Local Metrics| J[MRR, nDCG, Precision@K]
    I -->|LLM-as-a-Judge| K[Faithfulness, Relevancy]
Loading

Core 2: The Production Runtime (02_enterprise_production_rag)

Once validated, the architecture is hardened into a containerized microservices deployment designed to serve concurrent users with sub-second retrieval latency, robust state management, and continuous uptime.

  • πŸ”Œ Microservices Architecture: Decouples the application into two independent Docker containers: a lightweight Streamlit UI and a heavy FastAPI ML backend - ensuring clean separation of concerns, isolated fault tolerance, and independent scalability.
  • ⚑ Non-Blocking Async Backend: Preloads neural weights at boot via FastAPI lifespans. CPU-heavy retrieval tasks are offloaded to background threads (asyncio.to_thread), preventing event loop blocking.
  • 🧠 Stateful Query Rewriting: Implements LLM coreference resolution with a bounded sliding window memory (e.g., rewriting "What was their profit?" into "What was Apple's profit?") and automatic Ticker entity extraction, eliminating context-window bloat.
  • πŸ›‘οΈ Strict Structured Synthesis: Maps Gemini's Structured Outputs to Pydantic models, strictly enforcing evidence-backed answers, document citations, and an is_grounded flag to prevent hallucinations.
  • πŸ’Ύ Idempotent ETL Pipeline: Uses MD5 hashing for stable chunk IDs to prevent ChromaDB duplication, and injects explicit ticker metadata to improve vector search accuracy.
  • πŸ” Trace Attribution UI: Features an expandable Streamlit UI trace that exposes internal telemetryβ€”allowing users to audit latency, cross-encoder scores, rewritten queries, and retrieved snippets.
graph LR
    subgraph Frontend
        UI[Streamlit Dashboard]
    end

    subgraph Backend - FastAPI
        API[app.py\nEndpoints]
        CM[ChatSessionManager\nCoreference & History]
        RET[EnterpriseRetriever\nAsync Recall & Rerank]
        GEN[RAGGenerator\nGrounded Synthesis]
    end

    subgraph ML Models & Storage
        BGE[Local Bi-Encoder]
        XENC[Local Cross-Encoder]
        VDB[(ChromaDB)]
        GEMINI[Gemini API]
    end

    UI <-->|REST / JSON| API
    API --> CM
    CM <--> GEMINI
    API --> RET
    RET <--> BGE
    RET <--> VDB
    RET <--> XENC
    API --> GEN
    GEN <--> GEMINI
Loading

⚑ Architectural Upgrades & System Design

Feature Core 1 (Foundations) Core 2 (Production) Engineering Rationale
Retrieval Engine Synchronous Two-Stage Async Two-Stage + Metadata Filtering Both cores utilize BAAI bi-encoders and cross-encoders. Core 2 offloads the heavy neural network predictions to thread pools (asyncio.to_thread) to prevent blocking the FastAPI event loop, and adds exact-match entity filtering before reranking.
Data Processing In-memory LLM parsing Streamed ETL pipeline ingest_data.py processes SEC 10-K datasets in chunks with ticker injection to prevent CPU/RAM bottlenecking during vectorization.
State Management Stateless / Manual Conversational Session Store In-memory sliding window dictionary maps session_id to chat histories for multi-turn contextual awareness without memory leaks.
Response Synthesis Unstructured Strings Strict Pydantic Models Gemini 3.5 Beta API forces structured JSON outputs ensuring is_grounded boolean flags and exact source citations are tracked.
Latency Optimization Synchronous notebook cells Async FastAPI + Model Caching Models are cached locally via Docker volumes and loaded into global app state at server boot. Tasks are offloaded via asyncio.to_thread.

πŸ“ Repository Structure

enterprise-rag-evaluation
β”œβ”€β”€ 01_rag_foundations/
β”‚   β”œβ”€β”€ data/
β”‚   β”‚   β”œβ”€β”€ knowledge-base/       # Raw Markdown documents
β”‚   β”‚   └── tests.jsonl           # Golden dataset for RAGAS evaluation
β”‚   β”œβ”€β”€ notebooks/
β”‚   β”‚   └── RAG_foundations_Masterclass.ipynb  # Core educational notebook
β”‚   └── scripts/
β”‚       └── custom_eval.py        # Custom retrieval & generation metrics (nDCG, MRR, LLM-as-a-Judge)
β”‚
β”œβ”€β”€ 02_enterprise_production_rag/
β”‚   β”œβ”€β”€ api/
β”‚   β”‚   β”œβ”€β”€ app.py                # FastAPI asynchronous backend
β”‚   β”‚   β”œβ”€β”€ chat_manager.py       # Session memory & coreference resolution
β”‚   β”‚   └── generate.py           # Strictly grounded response generation
β”‚   β”œβ”€β”€ data/
β”‚   β”‚   └── chroma_sec_db/        # Persistent ChromaDB vector storage
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ ingest_data.py        # SEC-10K data chunking & embedding pipeline
β”‚   β”‚   └── retrieve.py           # Async Two-Stage Retriever (Recall + Rerank)
β”‚   β”œβ”€β”€ ui/
β”‚   β”‚   └── streamlit_app.py      # Interactive chat dashboard with trace attribution
β”‚   β”œβ”€β”€ evaluate/
β”‚   β”‚   └── evaluate_pipeline.py  # Automated RAGAS benchmarking pipeline
β”‚   └── assets/
β”‚       └── screenshots/          # πŸ“Έ Folder containing app running screenshots
β”‚           β”œβ”€β”€ ui_dashboard.png
β”‚           β”œβ”€β”€ engine_controls.png
β”‚           β”œβ”€β”€ ticker.png
β”‚           β”œβ”€β”€ connection_error.png
β”‚           β”œβ”€β”€ trace_attribution.png
β”‚           └── evaluation_metrics.png
β”‚
β”œβ”€β”€ docker-compose.yml            # Multi-container orchestration
β”œβ”€β”€ Dockerfile                    # Application environment specification
β”œβ”€β”€ .gitignore                    # Standard git ignores
β”œβ”€β”€ requirements.txt              # Manage Python project dependencies
β”œβ”€β”€ architecture_complete.jpg     # Dual-core system: laboratory to production runtime
└── README.md                     # Master architectural & deployment documentation


πŸ“Έ UI & Trace Attribution Gallery

All visual assets and execution screenshots for the dashboard, engine telemetry, and evaluation metrics are located in ./02_enterprise_production_rag/assets/screenshots/.

Dashboard Overview Trace Attribution Telemetry
UI Dashboard Trace Attribution
Engine Controls Evaluation Metrics
Engine Controls Evaluation Metrics

πŸ› οΈ Quickstart (Windows / Linux)

1. Clone & Configure

git clone https://github.com/Subhrajyoti8520/enterprise-rag-evaluation.git
cd enterprise-rag-evaluation
echo "GEMINI_API_KEY=your_api_key_here" > .env

2. Local Database Initialization (Recommended) To avoid Docker timeouts during the initial HuggingFace model downloads, build the vector database locally first.

# Create and activate a virtual environment
python -m venv venv
# Windows:
venv\Scripts\activate
# Mac/Linux:
source venv/bin/activate

# Install dependencies and run ingestion
pip install -r requirements.txt
python 02_enterprise_production_rag/src/ingest_data.py

(Note: By default, this runs in fast ingestion mode, sampling 500 records. Edit IngestionConfig in ingest_data.py to ingest the full dataset).

3. Launch the Microservices

docker-compose up --build -d

4. Run RAGAS Pipeline Evaluation (Optional)

python 02_enterprise_production_rag/evaluate/evaluate_pipeline.py

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

About

A production-style, two-stage Financial RAG architecture analyzing SEC 10-K reports. Features an evaluation sandbox (RAGAS) and an asynchronous, containerized FastAPI backend powered by Gemini Flash and ChromaDB.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages