Skip to content

Repository files navigation

🏒 AlphaLens

Financial Retrieval-Augmented Generation for SEC 10-K Filings

Build a searchable knowledge base from SEC annual reports using structure-aware parsing, semantic chunking, and multi-stage retrieval.

Live Demo

Python Gradio Gemini Qdrant Jina AI HuggingFace License: MIT

Stop settling for generic, hallucination-prone RAG pipelines. AlphaLens is a multi-stage Retrieval-Augmented Generation engine built specifically for SEC 10-K annual reports.


AlphaLens Dashboard


πŸ“Έ See it in Action

Knowledge Base Creation Financial Q&A
Download, parse, chunk, embed, and build the vector database with live pipeline logs. Ask financial questions and inspect the supporting source chunks used to generate every answer.

πŸš€ At a Glance

  • πŸ“„ Structure-aware SEC parsing
  • 🧠 Semantic chunking with text-fidelity validation
  • πŸ” Dual retrieval + query rewriting
  • 🎯 Cross-encoder reranking
  • πŸ“š Source-grounded answers

Why AlphaLens?

Most Retrieval-Augmented Generation systems are designed for blogs, documentation, and knowledge bases.

SEC filings are fundamentally different.

They contain hierarchical sections, financial tables, legal disclosures, and long contextual dependencies that should not be broken into arbitrary chunks.

AlphaLens was built specifically for these documents.

Note Chunk extraction is validated against the source filing. Final answers are still generated by an LLM and should always be verified using the cited source chunks.


✨ Key Features

Most RAG tutorials treat a 10-K like a generic PDF β€” blindly slicing it into 500-token chunks, destroying tables, splitting paragraphs mid-sentence, and hoping the LLM figures it out. That doesn't work for finance.

🧠 Domain-Aware Parsing Maps the document by SEC Items (1, 1A, 7...) instead of splitting by word count, keeping related financial context together.
πŸ”’ Verbatim Semantic Chunking An LLM identifies semantic boundaries, but 100% verbatim text preservation is strictly enforced β€” no dropped tables, no paraphrased risk factors.
🏷️ Auto-Enriched Metadata Every chunk gets an LLM-generated headline and summary, embedded alongside the text for stronger vector search.
πŸ” Advanced 6-Step Retrieval Query rewriting, dual retrieval, merging, and cross-encoder reranking find the exact financial detail in the haystack.
πŸ›‘οΈ Production-grade Resiliency Aggressive retry logic, rate-limit handling, and cold-start management β€” built to run reliably.

πŸ—οΈ Architecture

1. Ingestion Pipeline β€” building the knowledge base

Takes a company name, downloads the raw SEC filing, and converts it into a structured, searchable vector database.

flowchart TD
    A[User Input: Company & Year] --> B[Ticker Lookup β€” SEC EDGAR Mapping]
    B --> C[Download 10-K β€” MD Generator]
    C --> D["Stage 1: Domain-Aware Parsing
    β€’ Extract SEC Items (1, 1A, 7...)
    β€’ Smart-merge small items"]
    D --> E["Stage 2: LLM Semantic Chunking
    β€’ Gemini identifies boundaries
    β€’ 100% verbatim text enforcement
    β€’ Generate headline + summary"]
    E --> F["Stage 3: Embed & Upload
    β€’ Jina v4 embeddings
    β€’ Upsert to Qdrant Cloud"]
Loading

2. RAG Orchestrator β€” answering questions

Every question runs through a 6-step pipeline designed to maximize accuracy and minimize hallucination.

flowchart TD
    A[User Question] --> B["Step 1: Query Rewriting
    Optimize query for vector search"]
    B --> C["Steps 2–3: Dual Retrieval
    Search original + rewritten query
    Fetch from Qdrant Cloud"]
    C --> D["Step 4: Merge & Deduplicate
    Combine results from both scans"]
    D --> E["Step 5: Cross-Encoder Reranking
    Jina Reranker scores relevance
    Keep only top chunks"]
    E --> F["Step 6: Generation
    Build prompt (context + history)
    LLM generates final answer"]
Loading

πŸ› οΈ Tech Stack

Category Technology Purpose
Frontend Gradio Real-time UI for ingestion logs, chat, and source viewing
LLM Orchestration LiteLLM Unified interface for LLM providers (Gemini)
LLM (Chunking/RAG) Google Gemini 3.1 Flash Lite Semantic boundary detection & answer generation
Embeddings Jina Embeddings v4 High-accuracy financial text vectorization
Reranking Jina Reranker (Cross-Encoder) Precision scoring of retrieved chunks
Vector Database Qdrant Cloud Scalable, production-ready vector storage & search
Data Validation Pydantic Enforcing strict JSON schemas for LLM outputs
Resiliency Tenacity Retries & exponential backoff for API/DB timeouts
SEC Data edgartools Downloading structured 10-K filings from EDGAR

Example Workflow

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Company Name β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
 Download 10-K
       β”‚
       β–Ό
 Structure-aware Parsing
       β”‚
       β–Ό
 Semantic Chunking
       β”‚
       β–Ό
 Jina Embeddings
       β”‚
       β–Ό
 Qdrant
       β”‚
       β–Ό
 User Question
       β”‚
       β–Ό
 Query Rewrite
       β”‚
       β–Ό
 Retrieval
       β”‚
       β–Ό
 Reranking
       β”‚
       β–Ό
 Final Answer

πŸ’‘ How To Use

  1. Build the knowledge base β€” type a company name (e.g. "jp morgan", "nvidia") and year (e.g. 2024), then click πŸš€ Build Knowledge Base. Watch the real-time logs as it processes the 10-K.
  2. Download the raw data β€” once processing finishes, a download button for the raw .md file appears.
  3. Chat β€” ask complex financial questions, e.g. "What are the primary risk factors related to regulatory changes?"
  4. Verify sources β€” the right panel shows the exact Item, headline, and text chunk the AI used. No blind trust.

βš™οΈ Getting Started

1. Clone the repository

git clone https://github.com/YOUR_USERNAME/AlphaLens.git
cd AlphaLens

2. Create a virtual environment & install dependencies

python -m venv .venv
# Windows
.venv\Scripts\activate
# Mac/Linux
source .venv/bin/activate

pip install -r requirements.txt

3. Set up environment variables

Create a .env file in the root directory:

# Gemini API (chunking & RAG LLM)
GEMINI_API_KEY=your_gemini_api_key_here

# Jina AI API (embeddings & reranking)
jina=your_jina_api_key_here

# Qdrant Cloud (vector database)
QDRANT_URL=https://your-cluster-url.aws.cloud.qdrant.io
QDRANT_API_KEY=your_qdrant_api_key_here

4. Launch the app

python app.py

Then open http://localhost:7860.


πŸ“‚ Project Structure

AlphaLens/
β”œβ”€β”€ app.py                  # Gradio UI (main entry point)
β”œβ”€β”€ master_pipeline.py      # Core ingestion orchestrator
β”œβ”€β”€ ticker_extractor.py     # Company name β†’ SEC ticker mapping
β”œβ”€β”€ MD_Generator.py         # Downloads raw 10-K markdown from SEC EDGAR
β”œβ”€β”€ requirements.txt
β”‚
β”œβ”€β”€ ingest/                 # Data processing pipeline
β”‚   β”œβ”€β”€ parsing.py          #   Stage 1: SEC item extraction & merging
β”‚   β”œβ”€β”€ stage_2_worker.py   #   Stage 2: LLM semantic chunking (Gemini)
β”‚   └── stage_3_embed.py    #   Stage 3: Jina embeddings + Qdrant upload
β”‚
β”œβ”€β”€ rag_orchestrator/       # 6-step RAG engine
β”‚   β”œβ”€β”€ pipeline.py         #   Connects all 6 steps
β”‚   β”œβ”€β”€ rewrite.py          #   Step 1: query rewriting
β”‚   β”œβ”€β”€ retriever.py        #   Steps 2–3: dual vector search
β”‚   β”œβ”€β”€ merger.py           #   Step 4: dedup & merge
β”‚   β”œβ”€β”€ reranker.py         #   Step 5: cross-encoder reranking
β”‚   β”œβ”€β”€ prompt_builder.py   #   Step 6a: context & history formatting
β”‚   β”œβ”€β”€ generator.py        #   Step 6b: final answer generation
β”‚   β”œβ”€β”€ schema.py           #   Pydantic data models
β”‚   └── config.py           #   LLM & Qdrant configuration
β”‚
β”œβ”€β”€ finance_db/             # Local database storage
β”œβ”€β”€ Knowledge-base/         # Raw 10-K markdown storage
└── stage_1_json/, stage_2_json/   # Pipeline caching for resilience

πŸ“„ License

This project is licensed under the MIT License β€” see the LICENSE file for details.

About

Production-grade Financial RAG for SEC 10-K filings with structure-aware parsing, semantic chunking, multi-stage retrieval, Jina reranking, and source-grounded answers

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages