Skip to content

About

A modular FastAPI system for multilingual video translation — extracts audio/frames, uses RAG-enhanced Whisper STT with visual context, refines transcription via LLM, translates with Google Translate, and reconstructs dubbed video. Supports local CPU & Google Colab GPU.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Video Translation System

A modular FastAPI-based video translation system that uses RAG-enhanced speech recognition to provide context-aware transcription, translation, and audio dubbing.

Features

  • Video Processing: Extract audio and frames from videos
  • RAG-Enhanced Transcription: Use visual context to improve speech recognition accuracy
  • Local Processing: Run Whisper and LLM models locally (CPU) or in Google Colab (GPU)
  • Multiple Languages: Support for multiple source and target languages
  • Modular Architecture: Easy to extend and customize
  • FastAPI Backend: RESTful API for video translation
  • Comprehensive Logging: Track all intermediate files and processing stages

Architecture

Video Input → Audio/Frame Extraction → Vector DB (Frame Embeddings) 
    → Semantic RAG Analysis → Whisper STT (with context)
    → LLM Refinement (sentence completion) → Google Translate
    → gTTS → Video Reconstruction

Note: The LLM is used for sentence completion/refinement (fixing broken Whisper output), NOT translation. Translation is handled by Google Translate for reliability.

Installation

Prerequisites

  • Python 3.9 or higher
  • ffmpeg installed on your system
  • (Optional) CUDA-capable GPU for faster processing

Install System Dependencies

# Ubuntu/Debian
sudo apt-get update
sudo apt-get install ffmpeg

# macOS
brew install ffmpeg

# Windows
# Download from https://ffmpeg.org/download.html

Install Python Dependencies

pip install -r requirements.txt

Configuration

  1. Copy the example environment file:
cp .env.example .env
  1. Edit .env to configure settings:
    • WHISPER_MODEL: Whisper model size (tiny, base, small, medium, large)
    • LLM_MODEL: Model for sentence refinement (default: google/flan-t5-large)
    • WHISPER_DEVICE: cpu or cuda
    • LLM_DEVICE: cpu or cuda

Usage

Option 1: FastAPI Server

  1. Start the server:
python main.py
  1. The API will be available at http://localhost:8000

  2. API Documentation: http://localhost:8000/docs

  3. Upload and translate a video:

import requests

with open('video.mp4', 'rb') as f:
    files = {'video': f}
    data = {
        'target_language': 'Spanish',
        'source_language': 'auto',
        'use_rag': True
    }
    response = requests.post('http://localhost:8000/api/v1/translate', 
                           files=files, data=data)

job_id = response.json()['job_id']
print(f"Job ID: {job_id}")
  1. Check status:
status = requests.get(f'http://localhost:8000/api/v1/status/{job_id}')
print(status.json())

Option 2: Direct Pipeline Usage

from services.pipeline import TranslationPipeline
from pathlib import Path

pipeline = TranslationPipeline()
result = pipeline.process(
    video_path=Path("video.mp4"),
    target_language="French",
    source_language="auto",
    use_rag=True
)

print(f"Final video: {result['files']['final_video']}")

Option 3: Google Colab (GPU Acceleration)

For faster processing of Whisper and LLM on GPU:

  1. Open the Colab notebook: colab_whisper_llm.ipynb
  2. Upload your files to Colab
  3. Run the cells to process with GPU acceleration
  4. Download the results

The notebook includes standalone functions that can run independently in Colab.

Output Files

For each translation job, the following files are generated in outputs/{job_id}/:

  1. frames/ - Extracted video frames
  2. vector_db/ - ChromaDB vector database
  3. original_audio.wav - Extracted original audio
  4. transcription.json - Full Whisper transcription with timing
  5. transcription.txt - Plain text transcription
  6. rag_context.json - Visual context from frames
  7. translation.json - Translated segments with timing
  8. translation.txt - Plain text translation
  9. translated_audio.mp3 - Generated speech in target language
  10. final_video.mp4 - Final video with dubbed audio
  11. manifest.json - Index of all generated files

Module Structure

├── modules/
│   ├── video_processor.py      # Video/audio extraction & reconstruction
│   ├── frame_embedder.py       # CLIP frame embedding generation
│   ├── vector_store.py         # ChromaDB integration
│   ├── semantic_rag.py         # Semantic RAG with self-pruning
│   ├── rag_context.py          # Visual context generation
│   ├── speech_to_text.py       # Whisper STT
│   ├── transcription_refiner.py # LLM sentence completion (Flan-T5)
│   ├── simple_translator.py    # Google Translate wrapper
│   └── text_to_speech.py       # gTTS synthesis
├── api/
│   └── endpoints.py            # FastAPI endpoints
├── models/
│   └── schemas.py              # Pydantic models
├── services/
│   └── pipeline.py             # Main orchestration
├── utils/
│   ├── logger.py               # Logging utilities
│   └── file_manager.py         # File management
├── config.py                   # Configuration
├── main.py                     # FastAPI app
└── requirements.txt            # Dependencies

API Endpoints

POST /api/v1/translate

Upload and translate a video

Parameters:

  • video (file): Video file
  • target_language (string): Target language (e.g., "Spanish", "French")
  • source_language (string): Source language or "auto"
  • use_rag (boolean): Enable RAG context

Response:

{
  "job_id": "uuid",
  "status": "pending",
  "message": "Translation job queued successfully"
}

GET /api/v1/status/{job_id}

Check translation status

Response:

{
  "job_id": "uuid",
  "status": "completed",
  "progress": "Completed",
  "files": [
    {
      "file_type": "final_video",
      "file_path": "/path/to/final_video.mp4",
      "exists": true
    }
  ]
}

GET /api/v1/download/{job_id}/{file_type}

Download a specific file

Parameters:

  • job_id: Job identifier
  • file_type: Type of file (e.g., "final_video", "transcription_json")

Customization

Using Different Models

Edit .env:

# Whisper model sizes: tiny, base, small, medium, large
WHISPER_MODEL=large

# LLM for sentence refinement (CPU-optimized options)
LLM_MODEL=google/flan-t5-large      # Default, fast on CPU
# LLM_MODEL=google/flan-t5-xl       # Better quality, needs 16GB+ RAM
# LLM_MODEL=meta-llama/Llama-3.2-1B-Instruct  # GPU recommended

RAG Toggle

RAG can be configured in two modes via .env:

# Testing mode (always use RAG regardless of confidence)
RAG_ENABLE_SELF_PRUNING=False

# Production mode (skip RAG if confidence < 0.3)
RAG_ENABLE_SELF_PRUNING=True

Disable RAG

For faster processing without visual context:

pipeline.process(
    video_path=video_path,
    target_language="Spanish",
    use_rag=False  # Disable RAG
)

Performance Tips

  1. Use GPU: Set WHISPER_DEVICE=cuda and LLM_DEVICE=cuda in .env
  2. Use Colab: For free GPU access, use the provided Colab notebook
  3. Smaller Models: Use tiny or base Whisper for faster processing
  4. Reduce FPS: Lower FRAME_EXTRACT_FPS for fewer frames to process

Troubleshooting

"ffmpeg not found"

Install ffmpeg using your system package manager.

Out of Memory

  • Use smaller models (tiny/base for Whisper)
  • Disable RAG with use_rag=False
  • Process on CPU instead of GPU
  • Use Colab with higher RAM

Slow Processing

  • Use GPU instead of CPU
  • Use smaller models
  • Reduce frame extraction FPS
  • Use Colab for GPU acceleration

Roadmap

Coming Soon: Voice Dubbing

A voice dubbing feature is planned that will clone the original speaker's voice for the translated audio. Details TBD.

License

MIT License

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

About

A modular FastAPI system for multilingual video translation — extracts audio/frames, uses RAG-enhanced Whisper STT with visual context, refines transcription via LLM, translates with Google Translate, and reconstructs dubbed video. Supports local CPU & Google Colab GPU.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages