Product outcome
Return a useful, resumable first result window inside the user's time and resource posture even when states are unusually expensive or a retained frontier grows quickly. The first fixed 500K window must not consume the whole campaign envelope before Inkcheck can make an allocation decision.
This is a post-0.6 product requirement. It must improve anytime value without weakening evidence, changing proof semantics, or pretending to predict search-space exhaustion.
Evidence
The checked three-family 0.6 promotion gate (benchmarks/results/long-tail-promotion-v2.json) found:
- Dog Ink Adventure explored about 244K states before the 14-minute time ceiling, with no resumable 500K base checkpoint.
- Heresy II explored about 400K states before the 4 GiB memory ceiling, also before a 500K base checkpoint.
- The Intercept completed 500K in about 272 seconds and could run independent child windows efficiently.
Only one of three families reached the allocation-policy phase. The problem is therefore first-window sizing before long-tail selection or stopping.
Contract
- Begin an adaptive campaign with a bounded deterministic probe smaller than the current fixed first window.
- Measure at least states per second, elapsed time, peak/retained memory growth, frontier or checkpoint size where available, and whether useful evidence appeared.
- Choose the next window size deterministically from the persisted observation ledger and declared resource posture. Geometric growth is allowed only after a complete, reopenable result/checkpoint window.
- Enforce explicit minimum and maximum window sizes, total campaign ceilings, protected search floors, and integer state accounting.
- Persist the chosen size, policy/version, factual inputs, forecast range, uncertainty, and reason in human and machine progress/results.
- Fail closed when telemetry is absent or incompatible. A fallback may be conservative, but must remain deterministic and explicit.
- Never infer completeness, an asymptote, or absence of later findings from throughput/resource measurements.
- Preserve exact source/config/seed identity and current critical-evidence semantics.
Acceptance criteria
- Replay the pinned Dog Ink Adventure, The Intercept, and Heresy II families under matched state, depth, time, memory, and disk ceilings.
- Dog and Heresy each produce at least one useful reopenable result window before consuming 50% of their declared time or memory envelope, or return a checked finding explaining why that is impossible.
- The Intercept does not regress runtime/assertion/goal/authored-knot/visible-outcome evidence against the fixed-window baseline at matched states; terminal diversity remains reported separately.
- The policy makes identical window decisions for identical source, config, seed, posture, and telemetry ledger.
- Tests cover tiny budgets, very slow states, steep memory growth, missing telemetry, partial terminal windows, and aggregate ceiling boundaries.
- Human progress explains the bargain in plain language; MCP output exposes the compact decision and drill-down report ID.
- Checked comparison includes both matched-state and matched-wall-clock views and leaves live promotion gated on evidence.
Non-goals
- Predicting total search-space size or completion time for non-exhaustive stories.
- Promoting shadow long-tail stopping as part of this issue.
- Using AI to choose budgets or interpret story content.
Product outcome
Return a useful, resumable first result window inside the user's time and resource posture even when states are unusually expensive or a retained frontier grows quickly. The first fixed 500K window must not consume the whole campaign envelope before Inkcheck can make an allocation decision.
This is a post-0.6 product requirement. It must improve anytime value without weakening evidence, changing proof semantics, or pretending to predict search-space exhaustion.
Evidence
The checked three-family 0.6 promotion gate (
benchmarks/results/long-tail-promotion-v2.json) found:Only one of three families reached the allocation-policy phase. The problem is therefore first-window sizing before long-tail selection or stopping.
Contract
Acceptance criteria
Non-goals