Companion code for Post-Training for LLMs: SFT, Preference Optimization, RLVR, Agentic RL, and Distillation.
One module per chapter, one Streamlit page per chapter, one MLflow experiment per chapter.
Book Preview:https://drive.google.com/file/d/1eA2sGYUc0uWgfLtKb3k0pFgAgTwEiZjM/view?usp=sharing
Book Link: https://shop.beacons.ai/aiengineeringinsider/90ef9c43-585d-4420-b214-c37ee7455064
python3 -m venv .venv && source .venv/bin/activate
make install # CPU / Apple Silicon: everything except vLLM and verl
make install-gpu # add the rollout and scaled-RL stack (CUDA required)
make test # 41 offline tests, no GPU, no network, no API key
make ui # http://localhost:8501Every dependency is pinned to an exact version. Post-training libraries change
trainer defaults across minor releases, and an unpinned trl turns a
reproducible run into a story about a run.
configs/ one YAML per experiment; every file extends base.yaml
data/ raw, interim, processed shards with manifests
src/posttrain/
config.py config inheritance, CLI overrides, content hashing
seeding.py deterministic reset, per-step rollout seeds
tracking.py MLflow-backed result store with a JSONL fallback
data/ contract, breaks, dedup, decontaminate, mixture, synth/
train/ sft, dpo, grpo, dapo entrypoints
rewards/ verifiers, reward models, rubric judges
envs/ LangGraph agent environments (Ch 9)
eval/ harness, judge, gate, collapse monitor
merge/ merging and distillation (Ch 10)
ui/ Streamlit multipage app, one page per chapter
scripts/ thin CLI wrappers; see `make help`
reports/ one markdown report per experiment
tests/ offline test suite
| Tier | Hardware | Model ceiling | Chapters fully runnable |
|---|---|---|---|
| T0 | Apple M-series, 24 to 48 GB | Qwen3-0.6B LoRA | 1, 2, 4, 5 (partial), 10 (partial) |
| T1 | Single 24 GB GPU or Colab A100 | Qwen3-1.7B LoRA | 1 to 7 |
| T2 | 1x H100 80 GB rented | Qwen3-8B LoRA, RLVR | 1 to 9 |
| T3 | 8x H100 node | 8B full FT, agentic RL | All |
Every lab has a T0 fallback. If a chapter's headline configuration needs an
H100, the reduced-scope variant is in the same configs/ directory with a
_t0 suffix and is referenced from the chapter text.
No experiment exists unless it has a config file, a report in reports/, and
an MLflow run ID. posttrain.config.config_hash is what makes that enforceable:
two runs with the same hash used the same configuration, and a hash absent from
MLflow was never recorded.
The Chapter 2 synthesis pipeline defaults to Gemini, but --teacher echo
substitutes a deterministic stub. Filters, dedup, decontamination, and the
manifest writer all run against it, which is how those stages stay covered in
CI. Chapter 1's UI page has an offline toggle that swaps in a ChatML stub
tokenizer, so the template diff and the loss-mask heatmap work on a plane.
MIT. Book text is copyright AI Engineering Insider.