A hands-on workshop by Ashish Ranjan Jha, author of Mastering PyTorch.
Train a small transformer model from scratch, optimize it for production throughput, and deploy it as a scalable inference API - all in 3 hours.
Using a GPT-style language model (~8M parameters) trained on WikiText-2 as the anchor project, you'll work through the engineering steps that separate a working model from a production-ready system:
| Module | What you'll do |
|---|---|
| 01 - Model + Training | Build a modern transformer, structure a reproducible training pipeline with eval, checkpointing, and logging |
| 02 - Speedups + Stability | Add mixed precision (AMP), optimize DataLoaders, profile bottlenecks, and fix NaNs/OOMs |
| 03 - Inference + Export | Batch inference, dynamic quantization, TorchScript & ONNX export, benchmark latency |
| 04 - Deploy | Wrap in FastAPI, containerize with Docker, deploy to Google Cloud Run (via Cloud Shell) |
The fastest way to get started. No local install required - just a Google account and a browser.
| Module | Colab link |
|---|---|
| 01 - Model + Training | |
| 02 - Speedups + Stability | |
| 03 - Inference + Export | |
| 04 - Deploy |
Steps:
- Click a Colab link above
- In Colab, go to Runtime → Change runtime type → T4 GPU
- Run the first cell - it clones the repo and installs dependencies automatically
- Run the remaining cells in order
Tip: Colab gives you a free NVIDIA T4 GPU. The training speedup and optimization demos (AMP,
torch.compile, profiler) show real gains on GPU.
If you prefer running locally (or need Docker for Module 04):
git clone https://github.com/arj7192/pytorch-production-workshop.git
cd pytorch-production-workshop
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
python setup_check.py # verify everything is installed
jupyter lab notebooks/ # open the first notebookNote:
pip installdownloads PyTorch (~2 GB). An NVIDIA GPU is recommended but not required - CPU works, just slower.
- A Google account (for Colab and Cloud Shell) - or Python 3.10+ if running locally
- Docker (for local Module 04 only - Cloud Shell has it pre-installed)
- GCP project with billing (for Cloud Run deployment in Module 04 - attendees without billing can follow the instructor demo)
No advanced PyTorch experience needed - if you've trained a model before, you're ready.
├── notebooks/ # Hands-on workshop notebooks (start here)
│ ├── 01_model_and_training.ipynb
│ ├── 02_training_speedups.ipynb
│ ├── 03_inference_and_export.ipynb
│ └── 04_deploy.ipynb
│
├── src/ # Production-quality Python modules
│ ├── model.py # Transformer model definition
│ ├── data.py # Dataset and DataLoader utilities
│ ├── train.py # Configurable training script (CLI)
│ ├── evaluate.py # Evaluation and metrics
│ ├── export.py # TorchScript / ONNX export
│ └── utils.py # Reproducibility, logging, checkpointing
│
├── serve/ # Inference microservice
│ ├── app.py # FastAPI server
│ ├── Dockerfile # Production container
│ └── deploy.sh # GCP Cloud Run deployment
│
├── configs/ # Training configurations
│ ├── default.yaml # Full training config
│ └── fast_debug.yaml # Quick smoke-test config
│
├── scripts/ # Standalone utilities
│ ├── benchmark_dataloader.py # DataLoader perf comparison
│ └── profile_training.py # Training profiler with Chrome trace
│
├── tests/ # Smoke tests
│ └── test_model.py
│
├── reference/ # Takeaway reference cards
│ ├── training_checklist.md # Production training checklist
│ ├── debugging_guide.md # NaN/OOM/instability fixes
│ └── deployment_checklist.md # Deployment readiness checklist
│
├── setup_check.py # Pre-workshop environment validator
└── requirements.txt
Quick orientation, repo walkthrough, and what "production-ready" means for this workshop.
- Define a small GPT-style transformer with modern architecture patterns
- Structure a clean training loop: DataLoader → forward → loss → backward → step
- Add evaluation, checkpointing, and reproducibility from the start
- Understand why each component matters for production reliability
- Speed: Mixed precision (AMP + GradScaler), DataLoader tuning (
num_workers,pin_memory,persistent_workers),torch.compile - Stability: Gradient clipping, NaN/Inf detection, learning rate warmup, OOM prevention
- Profiling: Use PyTorch Profiler to find actual bottlenecks (not guessed ones)
- Batched inference for throughput
- Dynamic quantization to cut model size and latency
- Export via TorchScript (
torch.jit.script) and ONNX (torch.onnx.export) - Benchmark: eager vs compiled vs quantized vs ONNX
- Wrap the model in a FastAPI inference service (Colab)
- Set up GCP project and upload checkpoint to Cloud Storage (Colab)
- Switch to Google Cloud Shell for Docker build + Cloud Run deploy
- Health checks, graceful startup, and what to monitor
- Free copy of Mastering PyTorch ebook
- Workshop recording for replay
- This complete code repository
- Reference checklists for training, debugging, and deployment
- Certificate of completion
| Problem | Fix |
|---|---|
pip install fails on torch |
Try pip install torch --index-url https://download.pytorch.org/whl/cpu first, then re-run pip install -r requirements.txt |
ModuleNotFoundError: No module named 'src' |
Make sure you're running notebooks from the notebooks/ directory (Jupyter Lab opened from there) |
num_workers > 0 crashes on macOS |
This is a known fork-safety issue on macOS - use num_workers=0 (the default). The notebooks handle this. |
torch.compile fails |
Expected on some platforms (macOS, older GPUs). The notebooks catch this gracefully and continue. |
| ONNX export warnings | Warnings about deprecated operators are harmless - check that the verification step prints PASS. |
| Docker build fails | Make sure you're building from the repo root: docker build -f serve/Dockerfile -t pytorch-workshop-api . |
Ashish Ranjan Jha is Co-Founder and CEO at Nativ, a localization platform helping businesses unlock new markets. He previously led ML teams at Revolut (fraud detection, KYC) and XYZ Reality (computer vision for construction), and worked as a data scientist at Tractable and Sentiance. He is the author of three books: Mastering PyTorch (Packt, 2 editions), Fight Fraud with Machine Learning (Manning), and The Supervised Learning Workshop (Packt). He holds a Master's from EPFL and is a UK Exceptional Talent visa holder.
MIT - use this code however you like.