A modular FastAPI-based video translation system that uses RAG-enhanced speech recognition to provide context-aware transcription, translation, and audio dubbing.
- Video Processing: Extract audio and frames from videos
- RAG-Enhanced Transcription: Use visual context to improve speech recognition accuracy
- Local Processing: Run Whisper and LLM models locally (CPU) or in Google Colab (GPU)
- Multiple Languages: Support for multiple source and target languages
- Modular Architecture: Easy to extend and customize
- FastAPI Backend: RESTful API for video translation
- Comprehensive Logging: Track all intermediate files and processing stages
Video Input → Audio/Frame Extraction → Vector DB (Frame Embeddings)
→ Semantic RAG Analysis → Whisper STT (with context)
→ LLM Refinement (sentence completion) → Google Translate
→ gTTS → Video Reconstruction
Note: The LLM is used for sentence completion/refinement (fixing broken Whisper output), NOT translation. Translation is handled by Google Translate for reliability.
- Python 3.9 or higher
- ffmpeg installed on your system
- (Optional) CUDA-capable GPU for faster processing
# Ubuntu/Debian
sudo apt-get update
sudo apt-get install ffmpeg
# macOS
brew install ffmpeg
# Windows
# Download from https://ffmpeg.org/download.htmlpip install -r requirements.txt- Copy the example environment file:
cp .env.example .env- Edit
.envto configure settings:WHISPER_MODEL: Whisper model size (tiny, base, small, medium, large)LLM_MODEL: Model for sentence refinement (default:google/flan-t5-large)WHISPER_DEVICE: cpu or cudaLLM_DEVICE: cpu or cuda
- Start the server:
python main.py-
The API will be available at
http://localhost:8000 -
API Documentation:
http://localhost:8000/docs -
Upload and translate a video:
import requests
with open('video.mp4', 'rb') as f:
files = {'video': f}
data = {
'target_language': 'Spanish',
'source_language': 'auto',
'use_rag': True
}
response = requests.post('http://localhost:8000/api/v1/translate',
files=files, data=data)
job_id = response.json()['job_id']
print(f"Job ID: {job_id}")- Check status:
status = requests.get(f'http://localhost:8000/api/v1/status/{job_id}')
print(status.json())from services.pipeline import TranslationPipeline
from pathlib import Path
pipeline = TranslationPipeline()
result = pipeline.process(
video_path=Path("video.mp4"),
target_language="French",
source_language="auto",
use_rag=True
)
print(f"Final video: {result['files']['final_video']}")For faster processing of Whisper and LLM on GPU:
- Open the Colab notebook:
colab_whisper_llm.ipynb - Upload your files to Colab
- Run the cells to process with GPU acceleration
- Download the results
The notebook includes standalone functions that can run independently in Colab.
For each translation job, the following files are generated in outputs/{job_id}/:
frames/- Extracted video framesvector_db/- ChromaDB vector databaseoriginal_audio.wav- Extracted original audiotranscription.json- Full Whisper transcription with timingtranscription.txt- Plain text transcriptionrag_context.json- Visual context from framestranslation.json- Translated segments with timingtranslation.txt- Plain text translationtranslated_audio.mp3- Generated speech in target languagefinal_video.mp4- Final video with dubbed audiomanifest.json- Index of all generated files
├── modules/
│ ├── video_processor.py # Video/audio extraction & reconstruction
│ ├── frame_embedder.py # CLIP frame embedding generation
│ ├── vector_store.py # ChromaDB integration
│ ├── semantic_rag.py # Semantic RAG with self-pruning
│ ├── rag_context.py # Visual context generation
│ ├── speech_to_text.py # Whisper STT
│ ├── transcription_refiner.py # LLM sentence completion (Flan-T5)
│ ├── simple_translator.py # Google Translate wrapper
│ └── text_to_speech.py # gTTS synthesis
├── api/
│ └── endpoints.py # FastAPI endpoints
├── models/
│ └── schemas.py # Pydantic models
├── services/
│ └── pipeline.py # Main orchestration
├── utils/
│ ├── logger.py # Logging utilities
│ └── file_manager.py # File management
├── config.py # Configuration
├── main.py # FastAPI app
└── requirements.txt # Dependencies
Upload and translate a video
Parameters:
video(file): Video filetarget_language(string): Target language (e.g., "Spanish", "French")source_language(string): Source language or "auto"use_rag(boolean): Enable RAG context
Response:
{
"job_id": "uuid",
"status": "pending",
"message": "Translation job queued successfully"
}Check translation status
Response:
{
"job_id": "uuid",
"status": "completed",
"progress": "Completed",
"files": [
{
"file_type": "final_video",
"file_path": "/path/to/final_video.mp4",
"exists": true
}
]
}Download a specific file
Parameters:
job_id: Job identifierfile_type: Type of file (e.g., "final_video", "transcription_json")
Edit .env:
# Whisper model sizes: tiny, base, small, medium, large
WHISPER_MODEL=large
# LLM for sentence refinement (CPU-optimized options)
LLM_MODEL=google/flan-t5-large # Default, fast on CPU
# LLM_MODEL=google/flan-t5-xl # Better quality, needs 16GB+ RAM
# LLM_MODEL=meta-llama/Llama-3.2-1B-Instruct # GPU recommendedRAG can be configured in two modes via .env:
# Testing mode (always use RAG regardless of confidence)
RAG_ENABLE_SELF_PRUNING=False
# Production mode (skip RAG if confidence < 0.3)
RAG_ENABLE_SELF_PRUNING=TrueFor faster processing without visual context:
pipeline.process(
video_path=video_path,
target_language="Spanish",
use_rag=False # Disable RAG
)- Use GPU: Set
WHISPER_DEVICE=cudaandLLM_DEVICE=cudain.env - Use Colab: For free GPU access, use the provided Colab notebook
- Smaller Models: Use
tinyorbaseWhisper for faster processing - Reduce FPS: Lower
FRAME_EXTRACT_FPSfor fewer frames to process
Install ffmpeg using your system package manager.
- Use smaller models (tiny/base for Whisper)
- Disable RAG with
use_rag=False - Process on CPU instead of GPU
- Use Colab with higher RAM
- Use GPU instead of CPU
- Use smaller models
- Reduce frame extraction FPS
- Use Colab for GPU acceleration
A voice dubbing feature is planned that will clone the original speaker's voice for the translated audio. Details TBD.
MIT License
Contributions are welcome! Please open an issue or submit a pull request.