A training script is autoMIL-compatible when it honors the 6 contract items below. The framework treats it as an opaque process; the contract is the seam between framework and consumer.
Any language, any ML library: if the process writes a conforming result.json
and exits cleanly, the experiment tree records it correctly. The framework
never reads model weights, loss curves, or intermediate checkpoints.
A training script must:
-
Read
automil/config.yaml(or honor a--configflag if exposed). The orchestrator runs the script from the experiment's working directory (a git worktree);automil/config.yamlis reachable at the relative path. Consumer configuration (hyperparameters, dataset paths, scoring formula documentation) all lives here. The framework does not inject config values as command-line args. -
Honor the framework-owned device masks. Never override or unset
CUDA_VISIBLE_DEVICES,ROCR_VISIBLE_DEVICES,HIP_VISIBLE_DEVICES, orGPU_DEVICE_ORDINAL. CUDA receives its physical ID inCUDA_VISIBLE_DEVICES. On ROCm/Linux,ROCR_VISIBLE_DEVICESselects the physical host GPU and the HIP/CUDA compatibility masks expose it as logical device 0. CPU execution receives empty masks. -
Branch on
AUTOMIL_ACCELERATOR. Its value iscpu,cuda, orrocm.AUTOMIL_GPU=0remains a backward-compatible logical slot, including on CPU, and is not evidence that a physical GPU exists. Accelerator consumers use logical device 0 after masking (torch.device("cuda:0")works for both PyTorch CUDA and ROCm builds); CPU consumers ignoreAUTOMIL_GPU. -
Exit cleanly on
SIGTERMwith partial output written to result.json. The orchestrator sends SIGTERM at the cap boundary (Phase 4 / D-115); scripts that ignore it lose work and corrupt the cell budget bookkeeping. Write a result.json with"partial": trueand"status": "budget_killed"before exiting. See the SIGTERM handling section below. -
Write
result.jsonto the working directory matchingautomil/schemas/result.schema.json. The framework validates this at ingestion; malformed payloads transition the node tocrashedwith the schema location in the error message. The minimum valid payload is{"primary_value": <float>}. All other fields are optional. -
Declared env vars are present at startup. The framework's
automil checkvalidatesautomil/config.yaml: env.requiredBEFORE submit; missing vars fail with a named error rather than crashing the training script deep in execution. If your script readsMY_DATASET_ROOT, declare it underenv.requiredin config.yaml andautomil checkcatches an absent value before the experiment is even submitted.
The shipped reference is examples/sklearn-iris/train.py (~75 lines). It
demonstrates contract items 1, 2, 4, 5; item 3 selects CPU and item 6 is empty.
Read it as the executable spec.
The script follows this structure:
- install SIGTERM handler (Pattern B below)
- load iris dataset
- train LogisticRegression
- compute accuracy and F1
- write
result.jsonwith{"primary_value": accuracy, "metrics": {...}, "status": "completed"}
To run it end-to-end: uv run automil submit --node iris_001 --files examples/sklearn-iris/train.py.
A pytorch consumer adds GPU mask handling; the rest mirrors sklearn-iris.
import json, os, signal, sys
import torch
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")
_state = {"completed": False, "loss": 0.0}
def _write_result(*, status: str, partial: bool):
payload = {"status": status, "primary_value": -_state["loss"], "partial": partial}
open("result.json", "w").write(json.dumps(payload))
def _on_sigterm(signum, frame):
_write_result(status="budget_killed" if not _state["completed"] else "completed",
partial=not _state["completed"])
sys.exit(0)
signal.signal(signal.SIGTERM, _on_sigterm)
# ... train loop ...
_state["completed"] = True
_write_result(status="completed", partial=False)Key points for pytorch consumers:
- Branch on
AUTOMIL_ACCELERATOR; for either CUDA or ROCm,device = torch.device("cuda:0")addresses the single masked accelerator. Never infer accelerator presence fromAUTOMIL_GPU, which is also0on CPU. _state["loss"]updates at each epoch; SIGTERM during training writes the best-so-far loss.sys.exit(0)from the SIGTERM handler signals a graceful flush; the daemon distinguishes exit code 0 from crash (non-zero) and from silent process death.- Move training loop logic between
signal.signal(...)and_write_result(status="completed").
Two patterns. Pick by fold count.
For consumers that train multiple folds and write fold_<i>_result.json
per fold, the framework provides an aggregator. Call
register_sigterm_flush() once at startup; the helper installs a SIGTERM
handler that aggregates completed-fold files into result.json.
from automil.runtime_helpers import register_sigterm_flush
def main():
register_sigterm_flush() # must be called in main thread, before DataLoader init
for fold_i in range(get_fold_count()):
# ... train fold ...
# write fold_{i}_result.jsonThe handler (in src/automil/runtime_helpers.py) calls aggregate_folds() from
automil.cells.reconcile, merges completed folds into a single result.json with
"partial": true, and calls sys.exit(0). The orchestrator daemon then records
status: executed with metadata.budget_killed = True.
Constraint: call register_sigterm_flush() in the main thread, before creating
any DataLoader or threading.Thread. Python's signal.signal() raises
ValueError if called from a non-main thread.
For single-shot consumers (no fold structure), install your own handler.
The handler closes over a _state dict updated as training progresses.
Idempotent: a late SIGTERM after _state["completed"] = True writes
status: completed instead of status: budget_killed.
examples/sklearn-iris/train.py uses Pattern B.
Regardless of pattern, the handler must:
- Write
result.jsonFIRST (before any cleanup). - Exit via
sys.exit(0)NOTsys.exit(130). Exit code 0 tells the daemon the process completed gracefully; code 130 (or non-zero) is treated as a crash. - Be idempotent: if
result.jsonalready exists (normal completion before SIGTERM), the handler should either skip the write or overwrite with the same payload.
The contract is automil/schemas/result.schema.json (Draft 2020-12).
Required: primary_value (number). Optional: metrics (dict of
str -> number; validation-only under the val-firewall), held_out (dict of
str -> number; sealed test metrics, quarantined from search and revealed only
via automil certify), status (one of completed, crash, budget_killed,
cancelled), elapsed_seconds, peak_vram_mb, fold_results, partial.
additionalProperties: true means consumers may extend.
Two ingest rules sit on top of the schema. (a) Held-out-named keys inside
metrics (any key containing test, held_out, or heldout, or named
summary) fail the node closed with a val-firewall error: test values belong
only in held_out. (b) An optional validation_folds list of per-fold
validation projections, shaped
[{"fold_index": <int>, "metrics": {"val_...": <number>}, "primary_value": <number>}, ...],
lets the framework recompute primary_se (the cross-fold standard error
that feeds the Ladder keep-margin) from the per-fold primary values; when present,
the recomputed SE is preferred over a reported primary_se.
Minimum valid payload:
{"primary_value": 0.912}Full payload example (autobench / CCRCC consumer):
{
"primary_value": 0.845,
"metrics": {
"val_auc": 0.87,
"val_bacc": 0.81
},
"held_out": {
"test_auc": 0.87,
"test_bacc": 0.83
},
"status": "completed",
"elapsed_seconds": 4098,
"peak_vram_mb": 4500
}The framework validates result.json at ingestion via jsonschema.validate(...).
Malformed payloads transition the node to crashed with an error that
references this schema's location. Schema path in the error:
see automil/schemas/result.schema.json.
The primary_value field is the single scalar used by the experiment tree for
ranking and keep/discard (UCB scoring, Ladder margin). It is the
validation selection signal, never test. When a metrics block is
present the framework recomputes the primary_value from it at ingest (the
scoring.formula reducer, mean by default; CR-1b) and prefers the
recomputed value, logging any disagreement; a bare {"primary_value": ...}
payload is used as reported. Test metrics belong in
held_out instead, sealed away from search and revealed once via
automil certify. Higher is always better. For loss minimization, negate:
"primary_value": -val_loss.
Declare vars under automil/config.yaml: env.required:
env:
required:
- MY_DATASET_ROOT
- HF_HOME # optional cache; only declare if your script reads it
passthrough:
- MY_DATASET_ROOT
- HF_HOMEautomil check validates env.required at startup; any missing var fails
with Missing required env var: <name>; see automil/config.yaml: env.required.
env.passthrough controls forwarding from the orchestrator process into
experiment subprocesses. Use this list to opt in consumer-specific vars
(formerly auto-injected AUTOBENCH_ROOT is now consumer-declared via
this list per Phase 8 / DEC-01).
For CPU-only consumers with no external data dependencies (e.g. sklearn-iris where the dataset is bundled), both lists stay empty:
env:
required: []
passthrough: []Running automil check before submitting experiments verifies all required
vars are set. Integrate it into your workflow setup script or CI environment
validation to catch missing env vars before GPU time is wasted.
-
Writing result.json AFTER cleanup. If your training loop closes files, releases CUDA memory, then writes result.json, a SIGTERM during cleanup loses the partial result. Always write FIRST, then clean up. Pattern:
_write_result()at the top of the SIGTERM handler; cleanup (closing file handles, del model, torch.cuda.empty_cache()) only after the write returns. -
sys.exit(0)without writing partial. A SIGTERM handler that exits before the writer fires produces an empty archive directory; the orchestrator seesresult.jsonmissing and synthesises a crash status. The handler must call the writer explicitly. Common mistake: reusing aKeyboardInterrupthandler that just callssys.exit(0)without adapting it for SIGTERM semantics.
Both pitfalls have the same fix: install the SIGTERM handler early (top of main), before any resource acquisition, and always write result.json from inside the handler before exiting.
examples/sklearn-iris/train.py: shipped reference (DEC-02).src/automil/schemas/result.schema.json: result.json contract.src/automil/cli/check.py:automil checkenv.required validator.src/automil/runtime_helpers.py: multi-fold SIGTERM helper.docs/getting-started.md: project initialisation and first submit.