Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OCR Microservice

An OCR (Optical Character Recognition) microservice built with Python and FastAPI. This service accepts PDF files and images, extracts text using Tesseract OCR, and returns high-quality OCR results. It is designed to be modular, scalable, and easily integrated with applications like Arana Assistant.


Table of Contents


Project Overview

The OCR Microservice is designed to:

  • Accept PDF documents and images via an API.
  • Convert PDFs to images if necessary.
  • Detect the language of the document.
  • Perform OCR using the appropriate language model.
  • Return extracted text for further processing.

Architecture Details

Overview

The microservice consists of the following components:

  1. API Layer (FastAPI): Handles HTTP requests and responses.
  2. PDF to Image Conversion: Converts PDF files to images for OCR processing.
  3. Language Detection: Determines the document's language to select the appropriate Tesseract language model.
  4. OCR Processing (LlamaParse): Extracts text from images.

Component Details

  • API Layer: Utilizes FastAPI for handling file uploads and routing requests to the appropriate services.
  • PDF to Image Conversion: Uses pdf2image or PyMuPDF to convert PDF pages into images.
  • Language Detection: Implements langdetect or Tesseract's language detection on the first page to determine the document's language.
  • OCR Processing: Uses llamaparse

Setup Instructions

Prerequisites

  • Python 3.10+
  • Docker (Optional but recommended)
  • Tesseract OCR Engine
  • Poppler Utils (if using pdf2image for PDF conversion)

Clone the Repository

git clone https://github.com/yourusername/ocr-microservice.git
cd ocr-microservice

Virtual Environment Setup

python -m venv venv
source venv/bin/activate  # On Windows use `venv\Scripts\activate`

Install Dependencies

pip install -r requirements.txt

Install System Dependencies

On Ubuntu/Debian

sudo apt-get update
sudo apt-get install -y tesseract-ocr libtesseract-dev poppler-utils

On macOS (using Homebrew)

brew install tesseract poppler

Run the Application

Using Uvicorn

uvicorn app.main:app --reload

Using Docker

docker build -t ocr-microservice .
docker run -p 80:80 ocr-microservice

API Documentation

This project includes automatically generated API documentation accessible through two interactive interfaces:

  1. Swagger UI: Provides an interactive UI to test endpoints and view the API schema.
  2. ReDoc: Offers a clean, detailed view of the API documentation.

Accessing the Documentation

Once the FastAPI server is running, you can access the documentation at the following endpoints:

Replace localhost with your server's IP address if you're accessing it remotely.

Security Measures

Input Validation

  • File Type and Size Validation: Ensures only PDF and image files are accepted and that they are within acceptable size limits.
  • Input Sanitization: All inputs are sanitized to prevent injection attacks.

Authentication and Authorization

  • API Key Authentication: Implemented using FastAPI's security utilities. Clients must provide a valid API key to access the endpoints.
  • Rate Limiting: Limits the number of requests per client to prevent abuse.

Data Handling

  • Secure File Handling: Files are processed in memory or stored in secure temporary directories that are cleaned after processing.
  • Error Handling: Robust exception handling is in place to prevent sensitive information leakage.

Network Security

  • HTTPS Enforcement: It is recommended to run the service behind a secure gateway or proxy like Nginx with SSL certificates.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages