Surya Chandra PDF OCR is a small OCR pipeline for turning scanned PDFs into searchable PDFs. It is built around two OCR engines:
- Chandra extracts text.
- Surya provides page geometry for accurate invisible text placement.
- Hybrid mode (
chandra+surya) combines Chandra text with Surya geometry.
chandra+surya is the only production searchable-PDF mode. Neither Tesseract nor a Surya-only fallback is installed or selected by the production container.
The project is meant for people who have scanned documents and want a local, reproducible workflow that produces PDFs with selectable and searchable text. It is especially useful when text-only OCR output is not enough and the text layer needs to align with the scanned page.
Use this project if you need:
- Searchable PDFs from scanned PDF files.
- Local processing with Python, Docker, or a simple desktop GUI.
- Russian/English OCR hint by default (
rus+eng) for engines with explicit language selection. Surya and Chandra detect language automatically and ignore this hint. - Strict geometry behavior: hybrid output should fail loudly if the geometry sidecar is missing instead of silently producing a low-quality text layer.
This project is probably not the right fit if you need:
- A general document management system.
- A camera/scanner capture UI.
- CPU-only performance on large batches.
- A polished end-user desktop application with installers and automatic updates.
Recommended local setup:
- Windows with PowerShell or
cmd.exe. - Python 3.11 available through
py -3.11orpython. - NVIDIA GPU and current NVIDIA driver.
- Internet access on first setup to download Python packages and model weights.
uvis recommended but not required; the setup script falls back topip.
The bootstrap script intentionally creates two separate virtual environments because Surya and Chandra have different dependency stacks. Keeping them isolated prevents one engine from upgrading packages in a way that breaks the other engine.
.venv_surya.venv_chandra
Expected versions after a healthy setup:
.venv_surya:torch==2.11.0+<selected CUDA wheel>,torchvision==0.26.0+<selected CUDA wheel>,torchaudio==2.11.0+<selected CUDA wheel>,pillow>=10.2,<11.0..venv_chandra:torch==2.11.0+<selected CUDA wheel>,torchvision==0.26.0+<selected CUDA wheel>,torchaudio==2.11.0+<selected CUDA wheel>.
setup_dual_venv.cmd first attests host GPU index 0 against the UUID in
UNISCAN_GPU_DEVICE_ID, then selects the PyTorch CUDA wheel from that device
only. The UUID is host-local configuration and must not be committed. Any
missing, malformed, or mismatched device is a hard failure.
cu126for GPUs below compute capability7.5, for example GTX 1070 / Pascalsm_61.cu128for GPUs with compute capability7.5or newer.
You can override this manually with UNISCAN_TORCH_CUDA_FLAVOR=cu126 or UNISCAN_TORCH_CUDA_FLAVOR=cu128, but the automatic selection is the recommended path.
pillow is pinned below 11 only where Surya needs it. surya-ocr==0.17.1 requires pillow>=10.2,<11.0, while Chandra can currently run with a newer Pillow version. This is one of the reasons the project does not use one shared venv for both engines. The base uniscan package allows newer Pillow builds; the Surya venv is pinned separately by setup_dual_venv.cmd.
The setup script also downloads and verifies the model weights. Runtime is intentionally strict: the GUI and CLI require the local caches to exist and will fail instead of silently downloading weights or degrading output quality mid-run.
Model cache locations:
- Chandra weights:
.hf_cache_chandra, modeldatalab-to/chandra-ocr-2. - Surya weights:
.surya_cache, componentstext_detectionandtext_recognition. - Auxiliary HF/ModelScope caches:
.hf_cache_suryaand.modelscope_cache.
The Docker production OCR path was audit-accepted as complete on 2026-08-19.
Its image, storage layout, quality/performance evidence, accepted limitations,
and reopening conditions are recorded in
docs/audit/OCR_ACCEPTANCE_CLOSURE_2026-08-19.md.
The intended user path is: clone the repository, run setup_dual_venv.cmd, then run run_basic_gui.cmd, the CLI, or the durable HTTP API.
Completed baseline:
- Windows dual-venv setup for Surya and Chandra.
- CUDA PyTorch installation and verification in both venvs.
- Setup-time model weight downloads and cache verification.
- Strict runtime behavior when caches or geometry sidecars are missing.
- English README, menu labels, GUI copy, and web UI copy.
- Human-readable troubleshooting for expected setup warnings.
- Durable async HTTP job metadata with restart-visible
doneandinterruptedstates for external orchestrators. - Ten-page, content-addressed hybrid chunks with atomic manifests and verified resume after a process or host restart.
- A production-only
chandra+suryacontract; single-engine execution remains available only in explicit benchmark/diagnostic commands.
The HTTP job API is the cross-project integration boundary. For the standard request metadata, idempotency rules, OCR-owned queue semantics, and GPU coordination with the LLM orchestrator, see UniScan OCR Job Protocol and OCR Orchestrator And GPU Contract.
Last checked on 2026-04-21 with Python 3.11 on Windows. The setup script now auto-selects the PyTorch CUDA wheel from GPU compute capability. GTX 1070 / sm_61 requires cu126; cu128 is not compatible with that card because current PyTorch cu128 wheels start at sm_75.
These are Windows setup observations, not cross-platform dependency locks. A
separate Docker cu126 environment was observed on 2026-08-18 and differs in
several transitive versions. Its immutable image digest, complete external
package snapshots, and limitations are recorded in
constraints/observed/README.md. Do not copy
those observations into Windows setup or enforce them without clean builds.
Surya venv:
torch==2.11.0+cu126on GTX 1070 /sm_61, ortorch==2.11.0+cu128onsm_75+.torchvision==0.26.0+cu126on GTX 1070 /sm_61, ortorchvision==0.26.0+cu128onsm_75+.torchaudio==2.11.0+cu126on GTX 1070 /sm_61, ortorchaudio==2.11.0+cu128onsm_75+.surya-ocr==0.17.1pillow==10.4.0transformers==4.57.1huggingface-hub==0.36.2pypdfium2==4.30.0pypdf==6.10.2reportlab==4.4.10setuptools==70.2.0
Chandra venv:
torch==2.11.0+cu126on GTX 1070 /sm_61, ortorch==2.11.0+cu128onsm_75+.torchvision==0.26.0+cu126on GTX 1070 /sm_61, ortorchvision==0.26.0+cu128onsm_75+.torchaudio==2.11.0+cu126on GTX 1070 /sm_61, ortorchaudio==2.11.0+cu128onsm_75+.chandra-ocr==0.2.0pillow==12.1.1transformers==5.5.4huggingface-hub==1.11.0pypdfium2==4.30.0pypdf==6.10.2reportlab==4.4.10setuptools==70.2.0
git clone https://github.com/NixWrk/Surya_Chandra_PDF_OCR.git
cd Surya_Chandra_PDF_OCR
$env:UNISCAN_GPU_DEVICE_ID = (nvidia-smi --id=0 --query-gpu=uuid --format=csv,noheader,nounits).Trim()
# For Docker, also copy .env.example to ignored .env and store this UUID there.
.\setup_dual_venv.cmd
.\run_basic_gui.cmdThe first setup can take a while. The script installs both environments, installs CUDA builds of PyTorch, downloads the OCR model weights, and verifies that the required caches are ready before it exits successfully.
The setup script is safe to re-run. If the expected CUDA torch stack is already installed, it skips the forced torch reinstall and only verifies the environment and caches.
Runtime is GPU-only for the active OCR path:
cudais the default for Chandra and Surya.autois accepted only as a GPU metadata hint for schedulers; runtime still requires CUDA.cpumode is not a supported OCR execution path in this repository.
$env:UNISCAN_CHANDRA_DEVICE_POLICY = "cuda"
$env:UNISCAN_CHANDRA_REQUIRE_GPU = "1"
$env:UNISCAN_SURYA_TORCH_DEVICE = "cuda:0"
$env:UNISCAN_SURYA_REQUIRE_GPU = "1"
.\run_basic_gui.cmdAfter setup, the GUI lets you:
- Choose a PDF.
- Run the required
chandra+suryapipeline. - Optionally limit OCR to pages such as
1,3,5-8.
Before OCR starts, UniScan always creates an image-only copy of the source PDF and removes any existing text layer. Chandra/Surya run against that cleaned copy, and the final searchable PDF is built over the cleaned copy as well.
By default, the GUI overwrites the selected input PDF with the searchable version. Intermediate artifacts are written under outputs/.
Run this after setup if OCR is slow or if Chandra reports that CUDA is unavailable:
@'
import torch
print("torch:", torch.__version__)
print("cuda_available:", torch.cuda.is_available())
print("cuda_device_count:", torch.cuda.device_count())
print("cuda_device_0:", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "N/A")
'@ | .\.venv_chandra\Scripts\python.exe -Repeat for Surya:
@'
import torch
print("torch:", torch.__version__)
print("cuda_available:", torch.cuda.is_available())
print("cuda_device_count:", torch.cuda.device_count())
print("cuda_device_0:", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "N/A")
'@ | .\.venv_surya\Scripts\python.exe -A healthy GPU install should show a torch version containing +cu and cuda_available: True. For Surya, pillow should stay below 11.0 because surya-ocr==0.17.1 depends on that range. For Chandra, pillow 12.x is acceptable unless Chandra changes its own dependency constraints.
Use the Chandra environment for the main CLI:
.\.venv_chandra\Scripts\python.exe -m uniscan --helpBuild a searchable PDF in the default hybrid mode:
$env:UNISCAN_CHANDRA_DEVICE_POLICY = "cuda"
$env:UNISCAN_CHANDRA_REQUIRE_GPU = "1"
$env:UNISCAN_SURYA_TORCH_DEVICE = "cuda:0"
$env:UNISCAN_SURYA_REQUIRE_GPU = "1"
.\.venv_chandra\Scripts\python.exe -m uniscan searchable-pdf `
--pdf "D:\path\input.pdf" `
--mode chandra+surya `
--lang rus+eng `
--strictFull-document hybrid OCR is isolated into 10-page PDF chunks by default. Each chunk completes both Chandra recognition and Surya geometry before its searchable PDF is accepted. The service then merges the chunks and verifies contiguous page coverage, page count, order, and page dimensions. Every chunk input and output is published atomically and recorded with its SHA-256, byte size, page count, and serialized stage summary. The cache key includes the source PDF hash, effective OCR settings, and pipeline revision.
If a process, container, or host stops, a retry with the same PDF and settings reuses only completed chunks whose hash and PDF geometry still validate. A running, failed, missing, or modified chunk is processed again. HTTP jobs share a content-addressed cache outside the per-job work directory, so an idempotent retry can resume even though it receives a new job id. The cache is retained on failure and removed after the result PDF has been copied successfully.
This bounds Surya input size, releases engine subprocess VRAM between chunks, and
applies the engine timeout to one chunk instead of the whole document. Set
UNISCAN_HYBRID_CHUNK_PAGES to a different positive value to tune the tradeoff.
0 disables document chunking for diagnostics. Explicit pages= selections keep
the existing non-chunked path.
Useful commands:
.\.venv_chandra\Scripts\python.exe -m uniscan searchable-pdf --help
.\.venv_chandra\Scripts\python.exe -m uniscan benchmark-ocr --help
.\.venv_chandra\Scripts\python.exe -m uniscan prepare-compare-txt --help
.\.venv_chandra\Scripts\python.exe -m uniscan build-searchable-from-artifacts --help
.\.venv_chandra\Scripts\python.exe -m uniscan serve-http --helpGPU smoke/prewarm:
.\scripts\run_hybrid_gpu_smoke.ps1 -InputPdf "D:\path\input.pdf" -Pages 1The smoke script copies the input into outputs/gpu_hybrid_smoke, sets the
CUDA-only Chandra/Surya runtime variables, uses the persistent local model
caches, removes the original text layer before OCR, and runs chandra+surya
without modifying the original PDF.
Start the local web/API service:
.\.venv_chandra\Scripts\python.exe -m uniscan serve-http --host 127.0.0.1 --port 8000Open:
http://127.0.0.1:8000
Synchronous API:
curl -X POST "http://127.0.0.1:8000/searchable-pdf?mode=chandra+surya&lang=rus+eng&strict=1" \
-H "Content-Type: application/pdf" \
--data-binary "@input.pdf" \
-o output.searchable.pdfAsynchronous API:
curl -X POST "http://127.0.0.1:8000/api/jobs?mode=chandra+surya&lang=rus+eng&strict=1&filename=input.pdf" \
-H "Content-Type: application/pdf" \
-H "X-UniScan-Protocol: uniscan-ocr-job.v1" \
-H "X-Project-ID: zotero" \
-H "X-Service-ID: zotero-worker" \
-H "X-Task-ID: zotero:item:ABCD1234:ocr" \
-H "X-Request-ID: <uuid-per-http-attempt>" \
-H "X-Idempotency-Key: zotero:item:ABCD1234:ocr:v1" \
-H "X-Priority: batch" \
-H "X-GPU-Policy: cuda" \
-H "X-Estimated-VRAM-GB: 8" \
-H "X-Estimated-Pages: 42" \
--data-binary "@input.pdf"
curl "http://127.0.0.1:8000/api/jobs"
curl "http://127.0.0.1:8000/api/jobs/<job_id>"
curl "http://127.0.0.1:8000/api/jobs/<job_id>/metadata"
curl -L "http://127.0.0.1:8000/api/jobs/<job_id>/result" -o output.searchable.pdf
curl -X POST "http://127.0.0.1:8000/api/jobs/<job_id>/cancel"X-Idempotency-Key is the safe retry key. Repeating the exact same PDF and OCR
parameters with the same key returns the existing job with
idempotent_replay: true; reusing the key for different bytes or parameters
returns 409 Conflict. If the previous matching job ended as error,
interrupted, or cancelled, the same request creates a new job while the
failed job remains in history.
Queue model:
The async API accepts jobs from multiple projects, workers, containers, and
manual tools, but the OCR service runs exactly one OCR document at a time.
GET /api/jobs reports worker_concurrency: 1 with queue counts and active
jobs. Waiting jobs are ordered by priority (interactive, normal, batch,
low) and creation time; a running document is not preempted. This keeps
Chandra/Surya GPU use predictable while still allowing many callers to submit
work safely.
The synchronous /searchable-pdf endpoint uses the same internal OCR pipeline
lock as the async worker. A synchronous request may wait behind OCR work already
in progress; this avoids concurrent mutation of process-wide OCR environment
variables and keeps GPU use serialized.
Durability:
The async API keeps durable job metadata under UNISCAN_WORK_ROOT/jobs.
Each job directory contains input.pdf, metadata.json, events.jsonl, and,
after completion, result.pdf; jobs.sqlite3 keeps a service-owned write-side
index. Restart recovery is driven by the per-job metadata.json files, not by
using SQLite as the source of truth.
GET /api/jobs returns queue counts plus active and recent jobs,
GET /api/jobs/<job_id> returns the current job summary,
GET /api/jobs/<job_id>/metadata returns the persisted metadata file, and
GET /api/jobs/<job_id>/result downloads the completed searchable PDF.
Completed results remain discoverable after a service restart. Jobs that were
queued during a restart are requeued if input.pdf exists. Jobs that were
running during a restart are marked interrupted, so callers can retry from
their own durable source queue if needed.
Retention cleanup can be enabled with:
UNISCAN_JOB_CLEANUP_ON_START=1
UNISCAN_JOB_CLEANUP_INTERVAL_SECONDS=3600
UNISCAN_JOB_RETENTION_DAYS=30
UNISCAN_FAILED_JOB_RETENTION_DAYS=90GPU/LLM scheduling remains the responsibility of the external orchestrator. OCR
stores resource hints such as gpu_policy, estimated_vram_gb, and
estimated_pages, but does not reserve GPU slots itself. The OCR runtime itself
requires CUDA and fails loudly instead of falling back to CPU.
Existing text layers are always removed before OCR. delete_text_layer=0 and
delete_original_text_layer=false are rejected by the HTTP API.
OCR quality/performance tuning:
# Require Chandra's page-aware geometry sidecar instead of using page-1 fallback.
UNISCAN_CHANDRA_REQUIRE_SIDECAR=1
# OCR input render DPI used by the basic/web workflow (clamped to 72..400).
UNISCAN_OCR_RENDER_DPI=220
# Set to 0 to disable banded token alignment and use the full dynamic program.
UNISCAN_ALIGN_BAND=
# JPEG quality for image-only source PDFs; 0 restores lossless pixmap embedding.
UNISCAN_TEXTLESS_JPEG_QUALITY=85When UNISCAN_ALIGN_BAND is unset, alignment uses an automatic band of
max(64, 20% of the longer token sequence) and falls back to full alignment if
the band cannot form a complete path.
Docker is useful when you want a repeatable GPU runtime with persistent model caches.
docker compose build
docker compose up -dThe default Compose file is standalone: it uses the project-local default network and does not require a pre-created Docker network. For integration with an existing Zotero worker stack, opt in to the external shared network:
docker compose -f docker-compose.yml -f docker-compose.shared-network.yml up -dThe shared-network override expects the external network to exist. Create it once when setting up that integration (the default standalone deployment does not need this step):
docker network create zotero-automationThe Dockerfile defaults to TORCH_CUDA_FLAVOR=cu126 because that is the safer wheel for older GPUs such as GTX 1070. For newer sm_75+ GPUs, you can build with cu128:
docker compose build --build-arg TORCH_CUDA_FLAVOR=cu128Open:
http://localhost:8000
Stop:
docker compose downThe compose file mounts these local folders:
.hf_cache_chandrafor Chandra Hugging Face weights..hf_cache_suryafor Surya Hugging Face weights..surya_cachefor Surya model cache..modelscope_cachefor ModelScope cache.outputsfor work artifacts.PDFsfor optional input files.
Docker GPU requirements:
- NVIDIA driver on the host.
- Docker with GPU support.
- Docker Compose with exact UUID device reservations.
Quick GPU check:
docker run --rm --gpus "device=$env:UNISCAN_GPU_DEVICE_ID" `
nvidia/cuda:12.8.1-base-ubuntu22.04 `
nvidia-smi --id=0 --query-gpu=index,uuid,name --format=csv,noheaderThe project keeps heavyweight runtime files out of git. These folders are expected to be local:
.venv_surya.venv_chandra.hf_cache*.surya_cache.modelscope_cache.uv_cache.tmp_*outputs
If a model download is interrupted, deleting the incomplete cache for that engine and rerunning setup or OCR is often enough.
chandra+surya is the only mode accepted by searchable-pdf, the desktop GUI,
and both HTTP job endpoints. Chandra provides text and Surya provides geometry;
missing either engine is a hard failure. Large documents use the resumable
chunking contract described above.
The lower-level benchmark-ocr and artifact comparison commands retain
single-engine choices for diagnostics and quality evaluation. They are not
production fallbacks and are not exposed by the searchable-PDF service.
torch shows +cpu:
Run .\setup_dual_venv.cmd again. The current setup script requires CUDA PyTorch and fails if a CUDA build cannot be installed.
cuda_available: False:
Run the scoped Docker check above. UniScan refuses to start unless container GPU
index 0 maps to the permitted host UUID.
no kernel image is available for execution on the device:
The installed PyTorch CUDA wheel does not support your GPU architecture. For example, GTX 1070 is sm_61, while PyTorch cu128 wheels support sm_75+. Re-run setup_dual_venv.cmd; it auto-selects cu126 for older GPUs and verifies compatibility by running a tiny CUDA tensor before setup succeeds.
CUDA out of memory while loading Chandra:
The CUDA wheel is compatible, but Chandra's model does not fit into available VRAM. OCR does not fall back to CPU. Free VRAM, move the OCR service to a larger GPU, reduce concurrent GPU workloads outside OCR, or resubmit later through the external orchestrator.
pip's dependency resolver does not currently take into account all the packages that are installed:
This can appear during the forced CUDA PyTorch reinstall. In the Surya venv, PyTorch may temporarily pull pillow 12.x, which conflicts with surya-ocr==0.17.1; the setup script immediately pins Surya back to pillow>=10.2,<11.0 after the torch install. Treat the final verification as authoritative: Surya should end on pillow 10.4.0 or another <11.0 build, while Chandra can stay on pillow 12.x.
Repeated setup runs should not reinstall CUDA torch when the exact selected and verified CUDA stack is already present. If the script does reinstall torch, it means one of torch, torchvision, or torchaudio was missing, had a different version, or failed the CUDA tensor smoke-test.
WARNING: Ignoring invalid distribution ~orch:
This means a previous interrupted PyTorch uninstall left temporary ~* directories in the venv site-packages. The setup script removes these invalid pip leftovers at the start of each run before checking torch versions. If this warning appears during an already-running old setup attempt, let that run finish or stop it, then re-run the updated setup_dual_venv.cmd.
Warning: You are sending unauthenticated requests to the HF Hub:
This is a Hugging Face rate-limit warning, not a project failure. Chandra weights are large, about 10.6 GB for datalab-to/chandra-ocr-2, so unauthenticated downloads can be slow. If needed, set HF_TOKEN before running setup to use authenticated Hugging Face requests.
No module named uniscan:
Install the package into both venvs:
.\.venv_surya\Scripts\python.exe -m pip install -e .
.\.venv_chandra\Scripts\python.exe -m pip install -e .Chandra cache/weights preflight failed or Surya cache/weights preflight failed:
The project requires local model caches at runtime. Re-run setup to download and verify the weights:
.\setup_dual_venv.cmdExpected cache targets are .hf_cache_chandra for datalab-to/chandra-ocr-2 and .surya_cache for Surya text_detection plus text_recognition. If setup cannot download them, fix network/auth/firewall access first; the GUI will not perform lazy model downloads during OCR.
setup_dual_venv.cmd cannot find labels such as ensure_venv:
Make sure the file has Windows CRLF line endings. A normal git checkout on Windows should handle this.
src/uniscan/app- high-level OCR orchestration.src/uniscan/ocr- OCR engine adapters, benchmarks, geometry, and searchable PDF assembly.src/uniscan/ui/basic_ocr_gui.py- local Tkinter GUI.src/uniscan/web/service.py- local HTTP API and web UI.setup_dual_venv.cmd- Windows dual-venv bootstrap.run_basic_gui.cmd- Windows GUI launcher.Dockerfileanddocker-compose.yml- GPU container runtime.
Install dev dependencies into a venv and run tests:
python -m pip install -e ".[dev]"
python -m pytest -qThe test suite contains Russian OCR fixture text on purpose. That text is part of OCR behavior coverage, not UI copy.