Skip to content

Add logits archive, training cost records and make ac2 - #5

Merged
drewOrc merged 4 commits into
mainfrom
feat/logits-archive-ac2
Sep 23, 2026
Merged

drewOrc merged 4 commits into
mainfrom
feat/logits-archive-ac2

Conversation

@drewOrc

@drewOrc drewOrc commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Prepares PLAN step 2 (AC2). No full model is trained in this PR.

What changes

  • Logits archive (AC3). evaluate writes results/logits/<run>.npz: float32 validation and test logits, gold intent ids, and metadata (model_name, model_revision, seed, per_intent, train_rows, oos_train_rows, eval_per_intent, dataset_revision, label_space_sha256, git_commit, git_dirty, created_at). load_logits refuses an archive with a missing key, wrong type, misaligned shapes, NaN, out-of-range labels, or a different label space. Metrics in the results JSON are computed from the archive as written.
  • Manifest. Archives are gitignored; each one's SHA-256, size and row counts go into the committed results/logits-manifest.json. make verify-logits checks every entry (for copies downloaded from a Release).
  • Training cost (AC5 efficiency table). train_summary.json and the results JSON record parameter counts, wall time, steps, device and peak memory. MPS: sampled driver_allocated_memory (headline, includes allocator cache) and current_allocated_memory (lower bound); torch 2.14 has no MPS peak counter. CUDA: max_memory_allocated. CPU: ru_maxrss (process peak since start; bytes on macOS, KB on Linux, converted).
  • make ac2. Seeds 42/43/44 on configs/bert-base.yaml; skips seeds already archived and matching the manifest; reuses weights left by a crashed evaluation; FORCE=1 clears and reruns; deletes weights of seeds 43/44 after archiving; writes results/ac2.json; exit 1 on FAIL. Refuses records that are not the AC2 setup (other model, subsampled data, row counts off).
  • Results JSON moves to results/runs/ so report.py does not read ac2.json or the manifest as run records.

Verification

  • make lint, make test (110 passed), make smoke (now writes and re-reads a logits archive: 170 KB for 151 x 3 rows).
  • MPS probe checked with the tiny smoke model on an M4 (1 step).
  • Mutation checks, each turned tests red and was reverted: judge >= to > (1 failure), judge all to any (1), metadata missing keys tolerated (15), manifest hash never compared (2).

- archive.py: per-run .npz with float32 validation/test logits, gold
  intent ids and required metadata (model and revision, seed, k, OOS
  training rows, dataset revision, label-space hash, git commit, time).
  load_logits refuses missing keys, wrong types, shape mismatches, NaN,
  out-of-range labels and a different label space. Archives are
  gitignored; their SHA-256 goes into results/logits-manifest.json.
- efficiency.py: parameter counts, wall time, steps and peak memory
  (MPS sampled driver/tensor memory, CUDA allocator peak, CPU ru_maxrss
  with the macOS/Linux unit difference handled), recorded in
  train_summary.json and copied into the results JSON.
- ac2.py and `make ac2`: seeds 42/43/44, resumable, FORCE=1 reruns,
  deletes weights of seeds other than 42 after archiving, writes
  results/ac2.json and exits 1 on FAIL.
- Results JSON moves to results/runs/ so report.py does not read
  ac2.json or the manifest as run records.
- make smoke now writes and re-reads a logits archive.
Review follow-up for the AC2 runner and the logits archive.

- Resume only what the current config produced: results JSON and
  train_summary.json both record the config, and a seed counts as done
  only when the config matches (output paths aside) and the logits
  SHA-256 in the results JSON, the manifest and the file all agree.
  Weights trained with another config are cleared, not reused.
- FORCE clears every seed's weights, results JSON, archive and manifest
  entry before training anything; evaluate removes the old results JSON
  before writing a new archive.
- judge checks each record's run_name, config seed and training seed
  against its slot, and refuses two seeds sharing one archive.
- ac2.json and the printout list validation numbers per seed; README
  and docstring say tuning uses validation only.
- evaluate refuses a model whose id2label differs from intent_names.
- git_dirty ignores results/; train_summary records the git state.
- environment() records device, device name and deterministic mode.
- Tests for seed propagation, logit row alignment, git_dirty, the
  crash and FORCE scenarios, and weights kept after a failed evaluation.
- run_ac2 deletes results/ac2.json before doing anything, so a run with
  a changed config (or FORCE) that fails partway leaves no old PASS.
- ac2.json records the config (seed excluded) and each seed's logits
  SHA-256.
- Tests for: no verdict after a failed changed-config run, manifest sha
  disagreeing with the results JSON, an archive carrying another seed,
  evaluate removing the old results JSON before predicting, and the
  training summary seed with a seed other than 42.
- Document that FORCE deletes seed 42's kept weights and that only FORCE
  brings them back after clean-checkpoints; note the conservative config
  identity in RunConfig.identity().
@drewOrc
drewOrc merged commit b3a59e2 into main Sep 23, 2026
3 checks passed
@drewOrc
drewOrc deleted the feat/logits-archive-ac2 branch September 23, 2026 03:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant