Multi-GPU LLM training you can actually run for free. 20 self-contained notebooks that go from "first distributed run" to DPO/GRPO alignment, LLM-as-judge evaluation, and multimodal training — all sized for Kaggle's free 2×T4 (or Colab), with small open models (Qwen2.5-0.5B class).
Browse with one-click Colab links: sugeerth.github.io/gpu-training-notebooks
| Notebook | What you learn | Runs on |
|---|---|---|
| Simple_MultiGPU_Training | The simplest possible distributed fine-tuning run | Kaggle 2×T4 |
| Simple_MultiGPU_ActualTraining | Full-parameter training end-to-end, no LoRA | Kaggle 2×T4 |
| Simple_MultiGPU_Benchmark | 1 GPU vs 2 GPUs vs parallelism strategies, measured | Kaggle 2×T4 |
| Notebook | What you learn | Runs on |
|---|---|---|
| LoRA_QLoRA_FineTuning | LoRA vs QLoRA vs full fine-tuning, side by side | Kaggle 2×T4 |
| Modern_Full_FineTuning_PostTraining | Full-weight fine-tuning with the modern post-training stack | Kaggle 2×T4 |
| Distributed_Training_DeepSpeed | DPO + LoRA accelerated by DeepSpeed and Ray | Kaggle 2×T4 |
| Free_Distributed_SFT_DPO | The SFT → DPO pipeline entirely on free GPUs | Kaggle 2×T4 |
| Colab_Pro_A100_Max_Training | How big you can go on a single A100 40GB (7B) | Colab Pro A100 |
| Notebook | What you learn | Runs on |
|---|---|---|
| PostTraining_Core_Understanding | Watch a raw LM become an assistant, stage by stage | Kaggle 2×T4 |
| DPO_Training | DPO from scratch — preferences without a reward model | Kaggle 2×T4 |
| Simple_MultiGPU_DPO_RLHF | The full SFT → DPO → RLHF pipeline on a tiny model | Kaggle 2×T4 |
| Simple_MultiGPU_Alignment_Showdown | DPO vs KTO vs ORPO vs SimPO on the same data | Kaggle 2×T4 |
| GRPO_Reasoning_Training | GRPO — teaching step-by-step reasoning with pure RL | Kaggle 2×T4 |
| Notebook | What you learn | Runs on |
|---|---|---|
| LLM_as_Judge_Evaluation | Score model outputs with a judge LLM, self-contained | Kaggle 2×T4 |
| Multi_Trace_Agent_Evaluation | Evaluate agent reasoning across multiple traces, with retries | Kaggle 2×T4 |
| Notebook | What you learn | Runs on |
|---|---|---|
| Simple_MultiGPU_Multimodal | Fine-tune a vision-language model on image+text | Kaggle 2×T4 |
| Multimodal_LoRA_QLoRA_DPO | One VLM fine-tuned three ways: LoRA, QLoRA, DPO | Kaggle 2×T4 |
| Simple_MultiGPU_Diffusion | Teach Stable Diffusion a new concept with LoRA | Kaggle 2×T4 |
| Simple_MultiGPU_ImageClassification | Vision Transformer classification, distributed | Kaggle 2×T4 |
| Simple_MultiGPU_Audio | Fine-tune Whisper for speech recognition | Kaggle 2×T4 |
- Keep every line of code in notebook cells. Kaggle kernels cannot import local
.pyfiles — anything that matters is inlined, which is why the notebooks are self-contained. - Launch with
accelerate launch, notnotebook_launcher.notebook_launcherforks after CUDA is initialized and dies with cryptic CUDA re-init errors on Kaggle — 11 of these notebooks use theaccelerate launchpattern instead, none usenotebook_launcher.
dpo_train.py + dpo_results.json — a standalone-script variant of the DPO run and its saved
metrics, referenced by PostTraining_Core_Understanding.
requirements.txt covers local runs; on Kaggle/Colab each notebook installs what it needs in its
first cell.