Skip to content

Integrate quant: production-qualified training, scalable diagnostics and inference tooling - #2

Merged
asgersvenning merged 163 commits into
masterfrom
quant
Sep 11, 2026
Merged

asgersvenning merged 163 commits into
masterfrom
quant

Conversation

@asgersvenning

@asgersvenning asgersvenning commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Purpose

Integrate the quant branch into master after qualification and a completed 30-epoch production run on four B200 GPUs. The production configuration used floating-point FP16 AMP with model compilation, figures and W&B enabled. Fully quantized training with a demonstrated throughput and memory advantage remains an unfinished goal.

Changes

  • Add opt-in native CUDA INT8 Linear training, PTQ/QAT and deployment benchmarks, with checkpoint metadata and explicit unsupported-mode guards.
  • Add optimizer compilation and CUDA batch prefetch controls, cache/loader improvements, and portable compiled-model checkpoint handling.
  • Make large confusion matrices and dendrograms practical for training diagnostics: bounded dashboard images, compact SVG geometry, iterative rendering and retained local figures.
  • Preserve unseen labels in external inference, separate image discovery from training indexing, and expose --class-list candidate filtering without dropping images or ground truth.
  • Add reproducible UCloud qualification, restoration, staging, I/O calibration and pinned metrics-evaluation helpers.
  • Strengthen static/import checks, installed-wheel validation, benchmark reporting, and repository contribution/agent conventions.

This is a long-lived integration branch: 156 commits and 212 changed files at ad1fd10. Preserve the existing commits with a merge commit, rather than squashing or rewriting the development history.

Validation evidence

User-reported production run:

  • Four B200 GPUs; 30 epochs; 17h34m23s training-loop wall time; best checkpoint selected at epoch 30.
  • All recorded training/validation losses finite. Final validation species top-1 94.2353%, top-5 98.9686%; genus top-1 97.6693%, family top-1 99.5047%.
  • Figures and W&B functional with 12,632 species. The preceding four-rank restoration qualification reported model, optimizer, scheduler and scaler restored from the same checkpoint.
  • The production run used an earlier pinned branch revision. Subsequent inference and tooling changes have separate local checks; the production run does not validate every commit or optional mode.
  • Small paired experiments showed roughly 33% lower later-epoch training time and 34% lower peak allocated GPU memory with model compilation. Optimizer compilation, explicit prefetch and native INT8 did not establish additional useful gains in those experiments.

Local validation:

  • Ruff lint/format and import contracts passed.
  • Installed-wheel smoke passed: minimal imports, CLI help/resources, training, reload and prediction. Optional dendrogram dependencies were intentionally absent in the minimal environment.
  • 20 focused class-list/discovery tests passed at d9e57bd.
  • At ad1fd10, static/import checks and 29 quantization/CI-scope tests passed with oneDNN forced to AVX2. The original CI parity failure was reproduced under AVX2 and fixed without relaxing numerical tolerances. The default x86 PTQ/QAT activation range is now 0..127 in UINT8 storage; explicit reduce_range=False retains the older full-range recipe for validated VNNI targets.
  • CI now uses master pushes and feature PRs, avoiding duplicate quant push/PR runs and cancelling superseded CI.
  • Full local CPU-suite attempts were interrupted during the benchmark tests without reaching a final result; no full-suite pass is claimed. GitHub PR CI and dataset benchmarks are running. CUDA and slow/download-dependent tests require their intended environments.

Remaining limits

  • This provides evidence against substantial regressions in the exercised production configuration, not a controlled full-scale master-versus-quant equivalence study.
  • Native INT8 DDP and EMA are explicitly unsupported. Fully quantized training performance remains follow-up work.
  • Compilation graph breaks/recompile limits and DDP gradient-stride warnings remain performance follow-ups.
  • Cold shared-filesystem reads still require sufficient I/O concurrency or staging; the calibration helper measures the current workload/cache state.
  • Full in-domain held-out test inference/evaluation is still pending. External expert inference completed, but domain shift and vocabulary coverage must be retained when interpreting its metrics.
  • GPU/TensorRT dashboard workflows require appropriately configured runners and inputs; this PR does not provision them.

Please review and let required CI finish before merging. Auto-merge is not enabled.

Verify baseline and dataset provenance, calibrate on training samples, and evaluate reloaded integer artifacts on held-out synthetic, MNIST, and Blair data. Preserve per-level predictions and coverage for review.
Add a CUDA probe with INT8 weights, saved activations, and forward/backward GEMMs plus numerical and storage checks. Record workload-dependent speed and memory results, loader measurements, and the remaining optimizer, checkpoint, and convergence requirements.
…dates

Add isolated SGD and AdamW parameter dispatch, stochastic checkpoint continuation checks, and ordinary Linear coverage. Form backward scales in float32 to avoid FP16 underflow. Reject zero-gradient compiled probes and record corrected eager learning plus negative performance results; production QT integration remains ongoing.
Include implementation source and tensor metadata in the experimental parameter cache key. Add compiled-gradient regression coverage and replace the stale-cache diagnosis with corrected learning and performance measurements.
Replace unbounded queues and the daemon writer with bounded ordered futures and direct validated writes. Decode each sample once, propagate failures with reader cleanup, and cap automatic cache threads at 16 with a synchronous override. Add CPU/CUDA regressions and a reproducible comparison showing similar PNG throughput.
Move integer kernels into the optional runtime backend, prepare eligible Linear weights before optimizer creation, and report floating-point coverage. Preserve selected weight ties and restore quantized parameter types from checkpoint recipes. Cover AMP, masked heads, inference aliases, regularization, diagnostics, MuonAuxAdamW and training resume; validate minimal and CUDA installed wheels.
Keep AMP update gating and Muon counters outside compiled child updates. Initialize lazy state on the first real call, use tensor learning rates for scheduler updates, and save scalar rates for eager resume. Preserve Muon's explicit BF16 casts.

Expose --compile-optimizer independently of model compilation and record it in benchmark reports. Validate with static checks, the full suite (261 passed, 76 skipped, 1 expected EMA failure), and focused CUDA checks (81 passed). INT8 storage-kernel experiments remain separate.
Record dense MNIST quality, timing and whole-run versus later-phase allocation. Trace the 256 MiB Triton tuning buffer and document fake-tensor and compiler-cache failures that prevent shipping the experimental storage kernel.
Use explicit Triton row updates with a fake-tensor-safe custom operator, preserving integer storage, stochastic rounding, and tensor version invalidation. Tune TorchAO matrix kernels through a local CUDA-graph tuner without its 256 MiB cache-flush buffer.

The paired dense MNIST run reduces whole-training peak allocation from 236.30 to 154.00 MiB. Record that runtime and quality parity remain unproven. Keep the many-group outer optimizer fallback as a strict expected-failure regression.

Validation: static checks; full suite 261 passed, 85 skipped, 1 expected EMA failure; CUDA 44 passed and 1 expected compiler limitation; minimal installed-wheel smoke passed.
Handle the out-of-place primitive that Dynamo emits for tensor learning rates. This prevents optimizer-loop graph breaks and per-weight cache exhaustion without changing global compiler limits. Restore the many-group regression and extend it to AdamW.

Verify one-rounding update bounds, optimizer state, and unbiased sub-code updates with advancing CUDA RNG. Document the paired MNIST result: memory remains lower, but speed and quality parity are not established.

Validation: static checks; full suite 264 passed, 90 skipped, 1 expected EMA failure; CUDA regression set 108 passed plus focused arithmetic and RNG checks passed.
Tag repository batch indices so LazyDataset can defer redundant per-sample views. Preserve ordinary external batched fetches, default collation, dataset identity, spawned workers, shuffle RNG, and CUDA prefetch.

The small-image cache probe improves from 0.52M to 2.22M samples/s. Record larger-batch MNIST runs showing QT memory and warm-runtime gains, while retaining cold-start and accuracy-variation limits.

Validation: static checks; full suite 274 passed, 90 skipped, 1 expected EMA failure; CUDA loader and QT suite 108 passed.
Include matrix and update kernel sources in the training tensor's compiler fingerprint, preventing old opaque graphs from hiding operator changes during validation. Retain the opaque matrix path after the graph-visible experiment showed excessive startup cost and unstable real-data accuracy despite passing numerical checks.

Validation: static checks; full suite 274 passed, 90 skipped, 1 expected EMA failure; retained CUDA backend 52 passed.
Record explicit modes in benchmark reports and preserve existing defaults. Keep normalized weights beside their Linear consumer so embedding publication does not split the INT8 autograd boundary.

Validate portable checkpoints and eager resume across normalized heads and optimizer modes; compare embedding-loss gradients in FP32. Full CPU suite and focused CUDA regressions pass.
Add a separate three-seed MNIST profile with matched float and INT8 compilation modes, preserving ordinary compilation baselines. Retain failures, arguments, summaries and artifacts.

Record 31% lower INT8 peak allocation and shorter whole training calls, while explicitly retaining the mixed later-epoch speed result. Validate harness failure paths and both synthetic oracles.
Comment thread pyproject.toml

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR should bump version

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bumped the package to 0.2.0 in pyproject.toml and the lockfile (3cb478f). The isolated installed-wheel smoke passed.

Comment thread .gitignore

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding many specific cases is not necessarily wrong, but it is definitely an anti-pattern

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Consolidated the individual UCloud JSON exceptions into one directory-scoped pattern and removed a redundant temporary-directory pattern (3cb478f). Generated artifacts remain ignored.

Comment thread tests/README.md

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall, I am quite suspicious about rolling our own test collections, though it does just wrap pytest anyways.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no custom pytest collector: the shell wrapper selects the existing environment, sets CPU/headless defaults and forwards arguments to python -m pytest. Clarified this and documented direct pytest usage in tests/README.md (3cb478f).

Comment thread tests/README.md

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My impression is also that this PR suggest many new tests which are a bit overly pedantic, wasting both compute and time for minimal gain.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added --durations=15 to CI so expensive cases are visible, and documented the preference for observable-contract/regression coverage over implementation-detail assertions. I have not pruned tests indiscriminately: the affected 66-case suite took 41.50s locally (65 passed, one expected EMA failure), while unrelated benchmark-model tests made the broad local run substantially slower. Further consolidation remains a review target, rather than claiming this concern is fully resolved.

Comment thread mini_trainer/trainer.py
best_epoch = epoch
if output_dir is not None:
raw_model = model.module if hasattr(model, "module") else model
raw_model = model

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bit of a weird choice to use a while statement for this purpose, I do see that this won't cause an infinite loop, but there isn't any guarantee that this won't change in the future.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kept traversal because callers can supply nested compilation/distribution wrappers, but added identity-based cycle detection with a clear ValueError (3cb478f). A cyclic wrapper chain can no longer hang checkpoint saving. Compiled checkpoint/resume and CPU DDP tests passed.

Comment thread mini_trainer/train.py Outdated
log.info(f"Training restarted from checkpoint(s): {checkpoint}")

if compile_optimizer:
from mini_trainer.training.compilation import compile_optimizer as prepare_compiled_optimizer

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are we doing lazy imports?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This lazy import was unnecessary: the compilation module is already imported above. Moved the optimizer compiler alias into the normal module-level imports (3cb478f).

Comment thread mini_trainer/train.py Outdated
)
validate_type(nn_model, torch.nn.Module)
if quantized_training or getattr(nn_model, "_quantized_training_recipe", None):
from mini_trainer.modeling.quantized_training import prepare_quantized_training

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are we doing lazy imports?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moved this import to module scope too (3cb478f). The public quantized-training facade is safe to import; only its optional native backend remains lazy. Minimal installed-wheel imports and training passed without quantization extras.

Comment thread mini_trainer/predict.py Outdated
report.update(source=os.path.abspath(class_list), sha256=hashlib.sha256(contents).hexdigest())
with open(os.path.join(output_dir, "class_filter.json"), "w", encoding="utf-8") as handle:
json.dump(report, handle, indent=2)
print(

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should NOT use bare print statements like this.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replaced the bare print with the repository logger (3cb478f). Class-filter reporting remains in class_filter.json as well. The class-list inference integration tests passed.

Comment thread mini_trainer/integrations/gbif.py Outdated
return cached_result

with urlopen(req) as resp:
with urlopen(req, timeout=10) as resp:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are we adding a fixed timeout here? Seems like the responsibility for the issue this is supposed to solve is somewhere else, where?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the hardcoded 10-second policy (3cb478f). The helper now accepts an optional keyword-only socket timeout and otherwise uses urllib's process-wide default. A socket timeout cannot enforce a whole taxonomy lookup or job deadline; that policy belongs to the calling orchestration layer. The distinction is documented, and default/explicit forwarding was checked without network requests.

@asgersvenning asgersvenning left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is generally a PR with many useful features, but it does require a bit of cleanup before it can be pulled.

@asgersvenning asgersvenning self-assigned this Sep 11, 2026

Copy link
Copy Markdown
Owner Author

Review follow-up pushed through focused branches and explicit merges (head 400377d). Replied to each inline comment; threads remain open for your review.

  • Bumped package/lockfile to 0.2.0; simplified ignore patterns, imports and inference logging; guarded checkpoint wrapper cycles; removed the fixed GBIF timeout policy.
  • The failing setup test now creates its own repository and distinct master/quant pins, selects the test interpreter explicitly, and reports setup stderr instead of masking it with a missing-log error. Both standalone cases pass under Python 3.14.5; the GitHub matrix still needs to validate the update.
  • Validation: static checks; affected suite 65 passed / 1 expected EMA failure; minimal installed-wheel training/reload/inference passed. A broad local suite was interrupted in unrelated benchmark-model tests; no full-suite pass is claimed. CI now reports the 15 slowest tests.
  • Added a metadata-only PR statistics bot. It will activate after the workflow reaches master. No PR checkout/code execution, additional credentials or GPU infrastructure are used.

Current statistics, generated locally from the PR diff using the same categorization:

176 included files · +18,959 / −1,172 lines

Area Files Added Deleted
Core module 32 2,758 327
Tests 50 4,020 748
Benchmark tooling tests 22 3,915 0
Benchmarks 38 4,975 79
Developer tools and execution configs 25 2,909 3
CI and repository automation 4 340 9
Packaging and dependencies 2 17 3
User-facing README 1 17 1
Agent tooling 1 2 0
Other 1 6 2
Other Markdown (excluded from headline) 37 3,655 155

Root README.md is the initial explicit Markdown exception. Other feature documentation can be added to the workflow's featureMarkdown set. Counts represent the diff, not commit totals or a measure of effort.

@codecov-commenter

Copy link
Copy Markdown

Welcome to Codecov 🎉

Once you merge this PR into your default branch, you're all set! Codecov will compare coverage reports and display results in all future pull requests.

ℹ️ You can also turn on project coverage checks and project coverage reporting on Pull Request comment

Thanks for integrating Codecov - We've got you covered ☂️

@asgersvenning asgersvenning left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good now!

@asgersvenning
asgersvenning merged commit 5551c81 into master Sep 11, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants