Use this protocol before publishing a number in the README or research report.
- corpus:
data/science/train.txtanddata/science/validation.txt - split: approximately 90%/10% by bytes at paragraph boundaries within each source book
- seed:
42 - optimizer: AdamW, learning rate
3e-4 - batch size:
8 - context length:
128 - evaluation: full validation split at every checkpoint
- reporting: validation loss, token perplexity, bits per byte, measured timing, and stored tensor bytes when relevant
- Change one variable at a time.
- Keep seed, token budget, model width/depth, and evaluation split fixed.
- Report software versions, backend, device, source hashes, and commit.
- Report mean and standard deviation across at least three seeds before claiming a reliable quality difference.
- Do not rank byte and BPE models by token perplexity alone. Token units differ.
- For cache tests, assert numerical equivalence before measuring latency.
- Warm up the device, time repeated runs, use the same sampling settings, and synchronize asynchronous backends before reading the clock.
| ID | Variable | Primary metric | Secondary metric |
|---|---|---|---|
| 01 | learned positions vs RoPE | validation loss | length generalization |
| 02 | framework MHA vs SDPA | tokens/sec | peak memory |
| 03 | no cache vs KV cache | decode latency | speedup vs context |
| 04 | model size | loss per training token | memory and throughput |
| 05 | full tuning vs LoRA | quality at fixed budget | trainable parameters |
| 06 | FP32 vs INT8 CPU | latency and size | validation loss |
For every run, save:
config, seed, git commit, corpus hash, device, torch version,
training/validation checkpoints, best checkpoint, evaluation JSON,
and every benchmark timing sample
The benchmark scripts accept --output so results can be stored as machine-readable JSON rather than copied from a terminal.