Admit one train_schedule variable and an active reserve - #133
Conversation
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 79b26068ed
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| def _gate_disjoint(ctx: _Context, bundle: Mapping[str, Any], ledger: IdentityLedger) -> dict[str, int]: | ||
| reserved = [ | ||
| record | ||
| for record in ledger.identities.values() | ||
| if derived_state(record) == "EVAL_RESERVE" | ||
| ] | ||
| reserved = ledger.active_reserve_records(ctx.experiment_id) | ||
| hashes = {record["normalized_text_sha256"] for record in reserved} |
There was a problem hiding this comment.
Keep every live reserve out of training
When a shared ledger contains a current experiment's valid active reserve plus an EVAL_RESERVE belonging to another experiment, training data that overlaps the foreign reserve now passes because this gate checks only active_reserve_records(ctx.experiment_id). The same gap applies to EVAL_BOUND identities once another active set satisfies the slice counts. This silently contaminates evaluation material; keep binding/count validation scoped to the active set, but check training overlap against every EVAL_RESERVE and EVAL_BOUND identity in the ledger.
Useful? React with 👍 / 👎.
| try: | ||
| base_bundle = schedule_bundle(baseline) | ||
| cand_bundle = schedule_bundle(candidate) | ||
| except ValueError as exc: |
There was a problem hiding this comment.
Bind the declared schedule into the environment hash
When preflight and launch reuse the same baseline/candidate JSON paths, changing the four schedule fields in those files and the matching process environment between calls leaves effective_environment_hash() unchanged because neither the schedule variables nor HLX_SCIENTIFIC_VARIABLE are in ENV_KEYS, and the files are represented only by their paths. Both calls can therefore pass these comparisons while the documented parity check falsely attests that the launch uses the preflighted schedule; include these process inputs or sealed-file content digests in the hash.
Useful? React with 👍 / 👎.
The intended arm leaves patience and minimum epochs unset on the control and sets both on the candidate. Main would count those four sealed keys separately. Declared train_schedule still admits them as one variable.
The declared train_schedule arm is one variable. The test checks that the normalized bundles differ and that admission passes.
Summary
train_schedule. That declaration normalizesHYPERLEX_TRAIN_EPOCHS,HYPERLEX_EARLY_STOP,HYPERLEX_EARLY_STOP_PATIENCE, andHYPERLEX_EARLY_STOP_MIN_EPOCHSinto one bundle. Both arms must setHLX_SELECT_METRIC=classify_macro_f1_nonnone. Any other difference, an undeclared schedule change, or a selection-metric change fails admission. Experiments that declareHLX_SELECT_METRICkeep the existing single-key mode. Declaring the variable does not register or authorize an experiment.EVAL_RESERVEbound only to the current experiment.evaluation_reserved=truestays historical evidence and is not cleared onEVAL_SPENT. Spent identities do not count, do not satisfy overlap, and cannot be reused. Slice counts and the binding identity count use the active set. The events digest still covers the full append-only log.scripts/shadow/hyperlexical/loop.pyis unchanged. This pull request does not seal an experiment, allocate a reserve, edit a private ledger, launch training, score, or move BEST.Validation
Local workflow-equivalent suite
HYPERLEX_OFFLINE=1 HYPERLEX_NO_RATE_LIMIT=1 PYTHONPATH=src python -m pytest -q:Focused admission contract and parity tests: 29 passed on Python 3.12.3.
scripts/shadow/hyperlexical/loop.pysha2561aa395081d7709be3844bf2568d12100d73d931ecbc51af43c3d6e01dba77e2a.Do not merge from this pull request alone.