- repository structure and local-first command-line configuration
- decoder-only Transformer forward pass
- causal masking and next-token loss
- deterministic tests for shapes and causality
- byte-level tokenizer and text dataset
- gradient accumulation
- checkpoint save/resume with optimizer state
- validation loss and perplexity
- first small English corpus experiment
- temperature, top-k, and top-p sampling
- KV cache with full-forward equivalence tests
- latency benchmark with JSON output
- learned positions and RoPE implementations
- matched learned-position versus RoPE experiment
- AdamW versus SGD
- context lengths 128/256/512
- model-size scaling
- LoRA building block
- portable weight-only INT8 comparison, including quality, storage, and latency
- full fine-tuning versus LoRA quality comparison
Every experiment should have a configuration, a fixed seed, recorded metrics, and a short written conclusion.