Skip to content

[long-run] Run the corrected long-run promotion gate #219

Description

@chaoz23

Parent: #215. Depends on #216, #217, and #218. Also depends on chaoz23/inkbench#3, chaoz23/inkbench#4, and chaoz23/inkbench#5; freeze chaoz23/inkbench#6 and chaoz23/inkbench#7 before execution.

Product outcome

Determine whether the candidate converts a matched creator-owned resource grant into more sustained actionable value—or stops materially earlier and more honestly—without sacrificing InkCheck's early discovery or witness integrity.

Preregistered protocol

  • Freeze exact InkBench/adapters/analysis, dependency lock, Node/V8, compiler, baseline/candidate CLI bytes, fixtures/manifests, machine posture, serial order, endpoints, exclusions, and stopping rules.
  • Reject mixed/legacy fingerprints and do not resume the old 60-minute matrix.
  • Primary arms: pinned b2651f5 one-core portfolio, candidate with only the long-run change, and simple systematic/coverage controls. Keep product auto-concurrency as a labeled sensitivity arm.
  • Core families: rare-prefix, novelty-honeypot, compound-needle, false-novelty, delayed-consequence, order-dependent, revisit-after-mutation, and deep-corridor.
  • Add multi-bug Intercept-20/Heresy-II-30 and clean public The Intercept, Heresy II, and Dog Ink Adventure cases.
  • Use at least six held-out structural variants × five paired search seeds per generated family; fixtures are independent clusters and seeds are nested.
  • Run fresh 20-minute cells first; run a fresh 60-minute matrix only after identity, witness, stop, adapter, and telemetry audits pass.

Acceptance criteria

  • Zero lost/replay-invalid critical evidence, hard OOMs, malformed partial reports, mixed fingerprints, silent limit violations, or duplicate logical cells.
  • No planted bug loss; one-sided 95% upper bound for paired time-to-first-discovery ratio is no worse than 1.10 per required family.
  • Candidate satisfies at least one preregistered material-value path: new critical evidence on two independent fixtures; at least 25% lower retained GiB-minutes/cost at equal critical yield; at least 2× longer useful operation plus 25% better category-specific efficiency; or at least 25% cost saved by an earlier honest partial stop with no common-horizon evidence loss.
  • Terminal-state multiplicity alone cannot pass.
  • Compact shared checkpoints before frontier artifacts dominate campaigns #156's 500K Intercept-under-512-MiB or matched-value gate and exact split equivalence remain true.
  • Adapt first result windows to observed throughput and resource growth #190's useful reopenable result before 50% of Dog/Heresy time or memory envelope remains true.
  • Launched-cell, fixed-grant-complete, and observed-anytime outcomes are separate; resource stops use competing-risk/common-horizon analysis rather than naive independent censoring.
  • Full machine-readable artifacts, negative results, and family-level paired estimates are published.

Non-goals

  • General algorithm-superiority claims.
  • Retrofitting the stopped legacy matrix.
  • Promoting a policy on terminal diversity alone.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions