Main interface: image upload, OpenCV visual preprocessing, Ranjana rendering, and Nepali transliteration output
Ranjana Rendering, Nepali transliteration, English translation & Roman (IAST) transliteration followed by OCR confidence metrics and character-level CTC decoding breakdown
Lipi Snap is a deep learning OCR pipeline designed to recognize and decode full words in Ranjana script. It processes visual inputs to seamlessly generate a multi-stage output:
- Devanagari transliteration (the recognized text)
- English transliteration (IAST) with Nepali-style schwa deletion
- English translation (the translated text)
Built on a CRNN + CTC architecture, trained on 241k+ synthetic images, the model achieves robust and reliable recognition on both seen and unseen vocabularies.
🚀 Live Demo: Hugging Face Spaces
Note: For the older character-level CharCNN, see
CharCNN/or the Archive Documentation.
- Word-Level Ranjana Recognition — Predicts entire words in Ranjana script in a single pass.
-
Optimized Convolutional Backbone — Employs
$3 \times 3$ kernels with 1px padding ($2,024,038$ parameters) for robust spatial feature extraction within the CRNN. -
Multi-GPU Training — PyTorch
nn.DataParallelfor 2× GPU parallel training. - Dynamic Data Augmentation — Random affine transforms, blurs, and noise applied on-the-fly during training.
- Cross-Platform Device Detection — Auto-selects CUDA / MPS (Apple Silicon) / CPU.
-
Translation & Transliteration Pipeline:
- Devanagari → Roman (IAST) via
indic-transliteration - Devanagari → English via
deep-translator(Google Translate, no API key required) - Devanagari numeral auto-conversion (०१२ → 012) before translation
- Devanagari → Roman (IAST) via
-
Neo-Minimal Streamlit UI:
- Responsive dual-panel layout (side-by-side on desktop, stacked on mobile)
- Dynamic Ranjana font rendering injected at runtime (Normal / Stylish)
- OpenCV visual pipeline preview (Raw vs. Processed input)
- Interactive random test image sampling from unseen test set
- CTC greedy decoding path visualization with highlighted valid tokens
- Character-level confidence breakdown bars
- OCR metrics cards (character count · avg confidence)
| Metric | Value |
|---|---|
| Vocabulary | 69 characters (Devanagari vowels, consonants, marks, digits) |
| Model Parameters | 2,024,038 |
| Validation Accuracy | 99.65% with Val Loss of 0.0033 (Epoch 75 / 99) |
| Test Accuracy (word) | 98.14% (6,468 unseen synthetic words) |
| Test Accuracy (char) | 99.69% (unseen synthetic words) |
Train/Val Set (241,366 samples)
- WER: 0.20% · CER: 0.05% · Exact Match: 99.80%
Unseen Test Set (6,468 unique words, both font styles)
- WER: 1.86% · CER: 0.31% · Exact Match: 98.14%
# 1. Create & activate virtual environment
python3 -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows
# 2. Install dependencies
pip install -r requirements.txt
# 3. Generate synthetic training data (see tools/ folder)
# 4. Train the model
python model/train_crnn.py
# 5. Evaluate (CER / WER / Accuracy)
python model/evaluation_script.py
# 6. Run the web app
streamlit run app_word_crnn.pyLipi-Snap/
├── app_word_crnn.py # Streamlit UI & inference pipeline
├── model/
│ ├── train_crnn.py # CRNN + CTC training script
│ ├── inference_crnn.py # Single-image CLI inference
│ ├── evaluation_script.py # CER / WER / Exact-Match evaluation
│ └── best_crnn.pth # Saved model weights
├── data/
│ ├── synthetic_words/ # Training images + labels.csv
│ ├── test_synthetic_words/ # Test images + labels.csv
│ ├── nepali_words.txt # Source words for synthetic data generation
│ └── new_test_words.txt # unseen test words
├── tools/ # Data generation scripts
├── font/ # Ranjana display & image rendering fonts (.otf / .ttf)
├── CharCNN/ # Archived character-level model
├── requirements.txt
├── .gitignore
├── CharCNN_archive.md # Archived documentation for CharCNN
└── README.md
Ranjana fonts work by substituting Devanagari Unicode codepoints with Ranjana glyphs — so Ranjana looks different, but the underlying characters are Devanagari. The CRNN model learns to predict Devanagari Unicode indices directly from the Ranjana visual features, and Greedy CTC decoding collapses the raw frame-level output sequence into the final Devanagari string.
The pipeline then optionally:
- Applies
indic-transliteration(IAST) on Nepali Transliteration (Devanagari) output for English Transliteration - Runs Google Translate using
deep-translator(free) on the Devanagari output (additionally handling Devanagari numeral to English digit conversion explicitly)
- Training: 241,366 synthetic images, 2 Ranjana fonts, 90:10 train/val split with on-the-fly augmentation
- Test: 6,468 completely unseen words across both font styles
- Vocabulary: 69 Devanagari characters (
ँ ं ः अ आ … ् ० … ९) + 1<blank>CTC token
Word lists sourced from Brihat Sabdakosh, Nepal Bhasa, Wikipedia scrapes, and custom curation.
- Ranjana fonts sourced via Callijatra
- CRNN + CTC architecture inspired by "Nepal Script Text Recognition using CRNN CTC Architecture"
- Real-world testing using some Ranjana artwork screenshots from Calligraphy Nepal, Callijatra (Instagram) & more online sources.