Skip to content

Repository files navigation

Enterprise RAG & Structured Data Extraction Pipeline

CI Pipeline

Production-ready enterprise RAG and structured data extraction pipeline built with Python 3.11, FastAPI, Pydantic, and Instructor. Features self-correcting extraction graph loops, vector store search, and automated evaluation quality gates enforced in CI.


System Architecture

                               ┌───────────────────────────┐
                               │   FastAPI Microservice    │
                               │  (/extract, /graph, etc)  │
                               └─────────────┬─────────────┘
                                             │
                       ┌─────────────────────┴─────────────────────┐
                       │                                           │
                       ▼                                           ▼
          ┌─────────────────────────┐                 ┌─────────────────────────┐
          │ Vector Ingestion & RAG  │                 │ Extraction Service      │
          │ (Embeddings & Retrieval)│                 │ (Instructor + Pydantic) │
          └─────────────────────────┘                 └────────────┬────────────┘
                                                                   │
                                                                   ▼
                                                      ┌─────────────────────────┐
                                                      │  Self-Correction Graph  │
                                                      │     Pipeline Loop       │
                                                      └────────────┬────────────┘
                                                                   │
                                              ┌────────────────────┴────────────────────┐
                                              ▼                                         ▼
                                     [ Validation OK ]                         [ Validation Error ]
                                              │                                         │
                                              ▼                                         ▼
                                   ┌────────────────────┐                   ┌───────────────────────┐
                                   │ Return Validated   │                   │ Retry Extraction with │
                                   │ Structured Schema  │                   │ Error Feedback Context│
                                   └────────────────────┘                   └───────────────────────┘

Architecture Components

  • API Layer (FastAPI): Exposes asynchronous REST endpoints for raw document ingestion, vector similarity search, and structured data extraction.
  • Extraction Service (Instructor & OpenAI): Maps unstructured text into strictly typed Pydantic models, such as InvoiceSchema or PurchaseOrderSchema.
  • Graph Pipeline (ExtractionGraphPipeline): Orchestrates the extraction workflow. If extraction fails schema validation, the pipeline captures the errors and routes the request back to the extraction node with feedback context for self-correction, up to a configurable maximum retry limit.
  • Vector Store & Ingestion: Handles document chunking, embedding generation, and similarity searching for context retrieval.
  • Automated Evaluation Harness (scripts/evaluate.py): Runs offline validation against test datasets to guarantee schema accuracy metrics meet project targets (>85%) before merging code.

Key Features

  • Type-Safe Structured Extraction: Powered by Pydantic and Instructor for schema-validated field extraction.
  • Graph-Based Self-Correction Loop: Automatic error-feedback iteration logic that retries failed field extractions using previous validation context.
  • Vector Store & Retrieval: Embedded document ingestion and similarity search services for retrieval-augmented workflows.
  • Automated CI Evaluation Harness: Threshold-gated evaluation (make eval) integrated into GitHub Actions to block code merges if extraction accuracy falls below targets (e.g., <85%).
  • FastAPI Microservice Interface: REST endpoints for document ingestion, structured extraction, and graph execution.
  • Containerized Deployment: Docker Compose support for local development and containerized production execution.

Tech Stack

Category Technology
Language Python 3.11
Package Manager uv
Core Frameworks FastAPI, Pydantic v2, Instructor
Testing & Quality Pytest, Pytest-Cov, Ruff, MyPy
CI/CD GitHub Actions, GNU Make

Getting Started

Prerequisites

  • Python 3.11+
  • uv (fast Python package installer)
  • Docker & Docker Compose (optional, for containerized execution)

Installation

1. Clone the repository

git clone https://github.com/thithikhine1506/Enterprise-RAG-Extraction.git
cd Enterprise-RAG-Extraction

2. Install dependencies

make install

3. Configure Environment Variables

Create a .env file in the project root:

OPENAI_API_KEY=your_openai_api_key_here

Usage

Running the API Server

Start the FastAPI application locally:

uv run uvicorn src.main:app --reload

Access the interactive API documentation at:

http://127.0.0.1:8000/docs

Running with Docker

Build and run the containerized service:

make run

To stop containers and clean up volumes:

make down

Testing & Quality Control

This repository enforces strict code coverage, type checking, and accuracy evaluation thresholds.

Command Action
make test Runs unit tests with coverage reporting via Pytest
make lint Checks code style with Ruff
make format Automatically fixes code style and formats with Ruff
make typecheck Validates static types using MyPy
make eval Runs the offline evaluation harness and checks the accuracy threshold
make check-all Runs the full test suite and evaluation harness together

License

Distributed under the MIT License. See LICENSE for more information.

About

Production-ready enterprise RAG and structured data extraction pipeline using LangGraph, Instructor, and Pydantic with automated CI evaluation metrics.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages