Build a searchable knowledge base from SEC annual reports using structure-aware parsing, semantic chunking, and multi-stage retrieval.
Stop settling for generic, hallucination-prone RAG pipelines. AlphaLens is a multi-stage Retrieval-Augmented Generation engine built specifically for SEC 10-K annual reports.
- π Structure-aware SEC parsing
- π§ Semantic chunking with text-fidelity validation
- π Dual retrieval + query rewriting
- π― Cross-encoder reranking
- π Source-grounded answers
Most Retrieval-Augmented Generation systems are designed for blogs, documentation, and knowledge bases.
SEC filings are fundamentally different.
They contain hierarchical sections, financial tables, legal disclosures, and long contextual dependencies that should not be broken into arbitrary chunks.
AlphaLens was built specifically for these documents.
Note Chunk extraction is validated against the source filing. Final answers are still generated by an LLM and should always be verified using the cited source chunks.
Most RAG tutorials treat a 10-K like a generic PDF β blindly slicing it into 500-token chunks, destroying tables, splitting paragraphs mid-sentence, and hoping the LLM figures it out. That doesn't work for finance.
| π§ Domain-Aware Parsing | Maps the document by SEC Items (1, 1A, 7...) instead of splitting by word count, keeping related financial context together. |
| π Verbatim Semantic Chunking | An LLM identifies semantic boundaries, but 100% verbatim text preservation is strictly enforced β no dropped tables, no paraphrased risk factors. |
| π·οΈ Auto-Enriched Metadata | Every chunk gets an LLM-generated headline and summary, embedded alongside the text for stronger vector search. |
| π Advanced 6-Step Retrieval | Query rewriting, dual retrieval, merging, and cross-encoder reranking find the exact financial detail in the haystack. |
| π‘οΈ Production-grade Resiliency | Aggressive retry logic, rate-limit handling, and cold-start management β built to run reliably. |
Takes a company name, downloads the raw SEC filing, and converts it into a structured, searchable vector database.
flowchart TD
A[User Input: Company & Year] --> B[Ticker Lookup β SEC EDGAR Mapping]
B --> C[Download 10-K β MD Generator]
C --> D["Stage 1: Domain-Aware Parsing
β’ Extract SEC Items (1, 1A, 7...)
β’ Smart-merge small items"]
D --> E["Stage 2: LLM Semantic Chunking
β’ Gemini identifies boundaries
β’ 100% verbatim text enforcement
β’ Generate headline + summary"]
E --> F["Stage 3: Embed & Upload
β’ Jina v4 embeddings
β’ Upsert to Qdrant Cloud"]
Every question runs through a 6-step pipeline designed to maximize accuracy and minimize hallucination.
flowchart TD
A[User Question] --> B["Step 1: Query Rewriting
Optimize query for vector search"]
B --> C["Steps 2β3: Dual Retrieval
Search original + rewritten query
Fetch from Qdrant Cloud"]
C --> D["Step 4: Merge & Deduplicate
Combine results from both scans"]
D --> E["Step 5: Cross-Encoder Reranking
Jina Reranker scores relevance
Keep only top chunks"]
E --> F["Step 6: Generation
Build prompt (context + history)
LLM generates final answer"]
| Category | Technology | Purpose |
|---|---|---|
| Frontend | Gradio | Real-time UI for ingestion logs, chat, and source viewing |
| LLM Orchestration | LiteLLM | Unified interface for LLM providers (Gemini) |
| LLM (Chunking/RAG) | Google Gemini 3.1 Flash Lite | Semantic boundary detection & answer generation |
| Embeddings | Jina Embeddings v4 | High-accuracy financial text vectorization |
| Reranking | Jina Reranker (Cross-Encoder) | Precision scoring of retrieved chunks |
| Vector Database | Qdrant Cloud | Scalable, production-ready vector storage & search |
| Data Validation | Pydantic | Enforcing strict JSON schemas for LLM outputs |
| Resiliency | Tenacity | Retries & exponential backoff for API/DB timeouts |
| SEC Data | edgartools | Downloading structured 10-K filings from EDGAR |
ββββββββββββββββ
β Company Name β
ββββββββ¬ββββββββ
β
βΌ
Download 10-K
β
βΌ
Structure-aware Parsing
β
βΌ
Semantic Chunking
β
βΌ
Jina Embeddings
β
βΌ
Qdrant
β
βΌ
User Question
β
βΌ
Query Rewrite
β
βΌ
Retrieval
β
βΌ
Reranking
β
βΌ
Final Answer
- Build the knowledge base β type a company name (e.g. "jp morgan", "nvidia") and year (e.g. 2024), then click π Build Knowledge Base. Watch the real-time logs as it processes the 10-K.
- Download the raw data β once processing finishes, a download button for the raw
.mdfile appears. - Chat β ask complex financial questions, e.g. "What are the primary risk factors related to regulatory changes?"
- Verify sources β the right panel shows the exact Item, headline, and text chunk the AI used. No blind trust.
1. Clone the repository
git clone https://github.com/YOUR_USERNAME/AlphaLens.git
cd AlphaLens2. Create a virtual environment & install dependencies
python -m venv .venv
# Windows
.venv\Scripts\activate
# Mac/Linux
source .venv/bin/activate
pip install -r requirements.txt3. Set up environment variables
Create a .env file in the root directory:
# Gemini API (chunking & RAG LLM)
GEMINI_API_KEY=your_gemini_api_key_here
# Jina AI API (embeddings & reranking)
jina=your_jina_api_key_here
# Qdrant Cloud (vector database)
QDRANT_URL=https://your-cluster-url.aws.cloud.qdrant.io
QDRANT_API_KEY=your_qdrant_api_key_here4. Launch the app
python app.pyThen open http://localhost:7860.
AlphaLens/
βββ app.py # Gradio UI (main entry point)
βββ master_pipeline.py # Core ingestion orchestrator
βββ ticker_extractor.py # Company name β SEC ticker mapping
βββ MD_Generator.py # Downloads raw 10-K markdown from SEC EDGAR
βββ requirements.txt
β
βββ ingest/ # Data processing pipeline
β βββ parsing.py # Stage 1: SEC item extraction & merging
β βββ stage_2_worker.py # Stage 2: LLM semantic chunking (Gemini)
β βββ stage_3_embed.py # Stage 3: Jina embeddings + Qdrant upload
β
βββ rag_orchestrator/ # 6-step RAG engine
β βββ pipeline.py # Connects all 6 steps
β βββ rewrite.py # Step 1: query rewriting
β βββ retriever.py # Steps 2β3: dual vector search
β βββ merger.py # Step 4: dedup & merge
β βββ reranker.py # Step 5: cross-encoder reranking
β βββ prompt_builder.py # Step 6a: context & history formatting
β βββ generator.py # Step 6b: final answer generation
β βββ schema.py # Pydantic data models
β βββ config.py # LLM & Qdrant configuration
β
βββ finance_db/ # Local database storage
βββ Knowledge-base/ # Raw 10-K markdown storage
βββ stage_1_json/, stage_2_json/ # Pipeline caching for resilience
This project is licensed under the MIT License β see the LICENSE file for details.


