Skip to content

Latest commit

 

History

16 Commits

Folders and files

Repository files navigation

🌟 Lipi Snap — Ranjana Script Word Recognition

Live Demo on Hugging Face Spaces

Lipi Snap Main Interface — Ranjana Script and Nepali Transliteration
Main interface: image upload, OpenCV visual preprocessing, Ranjana rendering, and Nepali transliteration output

Lipi Snap — English Translation, IAST Transliteration, and Detailed Analysis
Ranjana Rendering, Nepali transliteration, English translation & Roman (IAST) transliteration followed by OCR confidence metrics and character-level CTC decoding breakdown


📖 Overview

Lipi Snap is a deep learning OCR pipeline designed to recognize and decode full words in Ranjana script. It processes visual inputs to seamlessly generate a multi-stage output:

  • Devanagari transliteration (the recognized text)
  • English transliteration (IAST) with Nepali-style schwa deletion
  • English translation (the translated text)

Built on a CRNN + CTC architecture, trained on 241k+ synthetic images, the model achieves robust and reliable recognition on both seen and unseen vocabularies.

🚀 Live Demo: Hugging Face Spaces

Note: For the older character-level CharCNN, see CharCNN/ or the Archive Documentation.


✨ Features

  • Word-Level Ranjana Recognition — Predicts entire words in Ranjana script in a single pass.
  • Optimized Convolutional Backbone — Employs $3 \times 3$ kernels with 1px padding ($2,024,038$ parameters) for robust spatial feature extraction within the CRNN.
  • Multi-GPU Training — PyTorch nn.DataParallel for 2× GPU parallel training.
  • Dynamic Data Augmentation — Random affine transforms, blurs, and noise applied on-the-fly during training.
  • Cross-Platform Device Detection — Auto-selects CUDA / MPS (Apple Silicon) / CPU.
  • Translation & Transliteration Pipeline:
    • Devanagari → Roman (IAST) via indic-transliteration
    • Devanagari → English via deep-translator (Google Translate, no API key required)
    • Devanagari numeral auto-conversion (०१२ → 012) before translation
  • Neo-Minimal Streamlit UI:
    • Responsive dual-panel layout (side-by-side on desktop, stacked on mobile)
    • Dynamic Ranjana font rendering injected at runtime (Normal / Stylish)
    • OpenCV visual pipeline preview (Raw vs. Processed input)
    • Interactive random test image sampling from unseen test set
    • CTC greedy decoding path visualization with highlighted valid tokens
    • Character-level confidence breakdown bars
    • OCR metrics cards (character count · avg confidence)

📈 Performance & Metrics

Metric Value
Vocabulary 69 characters (Devanagari vowels, consonants, marks, digits)
Model Parameters 2,024,038
Validation Accuracy 99.65% with Val Loss of 0.0033 (Epoch 75 / 99)
Test Accuracy (word) 98.14% (6,468 unseen synthetic words)
Test Accuracy (char) 99.69% (unseen synthetic words)

📊 Detailed Evaluation (via evaluation_script.py, using jiwer)

Train/Val Set (241,366 samples)

  • WER: 0.20% · CER: 0.05% · Exact Match: 99.80%

Unseen Test Set (6,468 unique words, both font styles)

  • WER: 1.86% · CER: 0.31% · Exact Match: 98.14%

🚀 Quick Start

# 1. Create & activate virtual environment
python3 -m venv .venv
source .venv/bin/activate        # macOS / Linux
# .venv\Scripts\activate         # Windows

# 2. Install dependencies
pip install -r requirements.txt

# 3. Generate synthetic training data (see tools/ folder)

# 4. Train the model
python model/train_crnn.py

# 5. Evaluate (CER / WER / Accuracy)
python model/evaluation_script.py

# 6. Run the web app
streamlit run app_word_crnn.py

📁 Project Structure

Lipi-Snap/
├── app_word_crnn.py             # Streamlit UI & inference pipeline
├── model/
│   ├── train_crnn.py           # CRNN + CTC training script
│   ├── inference_crnn.py       # Single-image CLI inference
│   ├── evaluation_script.py    # CER / WER / Exact-Match evaluation
│   └── best_crnn.pth           # Saved model weights
├── data/
│   ├── synthetic_words/        # Training images + labels.csv
│   ├── test_synthetic_words/   # Test images + labels.csv
│   ├── nepali_words.txt        # Source words for synthetic data generation
│   └── new_test_words.txt      # unseen test words
├── tools/                      # Data generation scripts
├── font/                       # Ranjana display & image rendering fonts (.otf / .ttf)
├── CharCNN/                    # Archived character-level model
├── requirements.txt
├── .gitignore
├── CharCNN_archive.md         # Archived documentation for CharCNN
└── README.md

🧠 How It Works

Ranjana fonts work by substituting Devanagari Unicode codepoints with Ranjana glyphs — so Ranjana looks different, but the underlying characters are Devanagari. The CRNN model learns to predict Devanagari Unicode indices directly from the Ranjana visual features, and Greedy CTC decoding collapses the raw frame-level output sequence into the final Devanagari string.

The pipeline then optionally:

  1. Applies indic-transliteration (IAST) on Nepali Transliteration (Devanagari) output for English Transliteration
  2. Runs Google Translate using deep-translator (free) on the Devanagari output (additionally handling Devanagari numeral to English digit conversion explicitly)

📊 Dataset

  • Training: 241,366 synthetic images, 2 Ranjana fonts, 90:10 train/val split with on-the-fly augmentation
  • Test: 6,468 completely unseen words across both font styles
  • Vocabulary: 69 Devanagari characters (ँ ं ः अ आ … ् ० … ९) + 1 <blank> CTC token

Word lists sourced from Brihat Sabdakosh, Nepal Bhasa, Wikipedia scrapes, and custom curation.


🔖 Acknowledgments

About

Lipi Snap: Ranjana Lipi - An OCR (Transcription, Transliteration & Translation) model for the Ranjana Lipi.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages