Skip to content

Admit one train_schedule variable and an active reserve - #133

Merged
scrimshawlife-ctrl merged 7 commits into
mainfrom
cursor/admission-schedule-reserve-41af
Sep 27, 2026
Merged

scrimshawlife-ctrl merged 7 commits into
mainfrom
cursor/admission-schedule-reserve-41af

Conversation

@scrimshawlife-ctrl

Copy link
Copy Markdown
Owner

Summary

  • A preregistered experiment may declare exactly one composite scientific variable, train_schedule. That declaration normalizes HYPERLEX_TRAIN_EPOCHS, HYPERLEX_EARLY_STOP, HYPERLEX_EARLY_STOP_PATIENCE, and HYPERLEX_EARLY_STOP_MIN_EPOCHS into one bundle. Both arms must set HLX_SELECT_METRIC=classify_macro_f1_nonnone. Any other difference, an undeclared schedule change, or a selection-metric change fails admission. Experiments that declare HLX_SELECT_METRIC keep the existing single-key mode. Declaring the variable does not register or authorize an experiment.
  • Active reserve membership is the derived lifecycle EVAL_RESERVE bound only to the current experiment. evaluation_reserved=true stays historical evidence and is not cleared on EVAL_SPENT. Spent identities do not count, do not satisfy overlap, and cannot be reused. Slice counts and the binding identity count use the active set. The events digest still covers the full append-only log.
  • scripts/shadow/hyperlexical/loop.py is unchanged. This pull request does not seal an experiment, allocate a reserve, edit a private ledger, launch training, score, or move BEST.

Validation

Local workflow-equivalent suite HYPERLEX_OFFLINE=1 HYPERLEX_NO_RATE_LIMIT=1 PYTHONPATH=src python -m pytest -q:

  • Python 3.10.21: 1051 passed, 13 skipped, 16 subtests passed
  • Python 3.11.16: 1051 passed, 13 skipped, 16 subtests passed
  • Python 3.12.3: 1059 passed, 7 skipped, 16 subtests passed

Focused admission contract and parity tests: 29 passed on Python 3.12.3.

scripts/shadow/hyperlexical/loop.py sha256 1aa395081d7709be3844bf2568d12100d73d931ecbc51af43c3d6e01dba77e2a.

Do not merge from this pull request alone.

A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.
A preregistered train_schedule declaration compares the four schedule fields as a single scientific variable. Holdout counts use only current EVAL_RESERVE identities bound to the experiment, leaving spent evaluation_reserved flags and the append-only ledger unchanged.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 79b26068ed

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines 570 to 572
def _gate_disjoint(ctx: _Context, bundle: Mapping[str, Any], ledger: IdentityLedger) -> dict[str, int]:
reserved = [
record
for record in ledger.identities.values()
if derived_state(record) == "EVAL_RESERVE"
]
reserved = ledger.active_reserve_records(ctx.experiment_id)
hashes = {record["normalized_text_sha256"] for record in reserved}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep every live reserve out of training

When a shared ledger contains a current experiment's valid active reserve plus an EVAL_RESERVE belonging to another experiment, training data that overlaps the foreign reserve now passes because this gate checks only active_reserve_records(ctx.experiment_id). The same gap applies to EVAL_BOUND identities once another active set satisfies the slice counts. This silently contaminates evaluation material; keep binding/count validation scoped to the active set, but check training overlap against every EVAL_RESERVE and EVAL_BOUND identity in the ledger.

Useful? React with 👍 / 👎.

Comment on lines +700 to +703
try:
base_bundle = schedule_bundle(baseline)
cand_bundle = schedule_bundle(candidate)
except ValueError as exc:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bind the declared schedule into the environment hash

When preflight and launch reuse the same baseline/candidate JSON paths, changing the four schedule fields in those files and the matching process environment between calls leaves effective_environment_hash() unchanged because neither the schedule variables nor HLX_SCIENTIFIC_VARIABLE are in ENV_KEYS, and the files are represented only by their paths. Both calls can therefore pass these comparisons while the documented parity check falsely attests that the launch uses the preflighted schedule; include these process inputs or sealed-file content digests in the hash.

Useful? React with 👍 / 👎.

The intended arm leaves patience and minimum epochs unset on the control and sets both on the candidate. Main would count those four sealed keys separately. Declared train_schedule still admits them as one variable.
The declared train_schedule arm is one variable. The test checks that the normalized bundles differ and that admission passes.
@scrimshawlife-ctrl
scrimshawlife-ctrl merged commit fab0de0 into main Sep 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant