Skip to content

Adapt first result windows to observed throughput and resource growth #190

Description

@chaoz23

Product outcome

Return a useful, resumable first result window inside the user's time and resource posture even when states are unusually expensive or a retained frontier grows quickly. The first fixed 500K window must not consume the whole campaign envelope before Inkcheck can make an allocation decision.

This is a post-0.6 product requirement. It must improve anytime value without weakening evidence, changing proof semantics, or pretending to predict search-space exhaustion.

Evidence

The checked three-family 0.6 promotion gate (benchmarks/results/long-tail-promotion-v2.json) found:

  • Dog Ink Adventure explored about 244K states before the 14-minute time ceiling, with no resumable 500K base checkpoint.
  • Heresy II explored about 400K states before the 4 GiB memory ceiling, also before a 500K base checkpoint.
  • The Intercept completed 500K in about 272 seconds and could run independent child windows efficiently.

Only one of three families reached the allocation-policy phase. The problem is therefore first-window sizing before long-tail selection or stopping.

Contract

  1. Begin an adaptive campaign with a bounded deterministic probe smaller than the current fixed first window.
  2. Measure at least states per second, elapsed time, peak/retained memory growth, frontier or checkpoint size where available, and whether useful evidence appeared.
  3. Choose the next window size deterministically from the persisted observation ledger and declared resource posture. Geometric growth is allowed only after a complete, reopenable result/checkpoint window.
  4. Enforce explicit minimum and maximum window sizes, total campaign ceilings, protected search floors, and integer state accounting.
  5. Persist the chosen size, policy/version, factual inputs, forecast range, uncertainty, and reason in human and machine progress/results.
  6. Fail closed when telemetry is absent or incompatible. A fallback may be conservative, but must remain deterministic and explicit.
  7. Never infer completeness, an asymptote, or absence of later findings from throughput/resource measurements.
  8. Preserve exact source/config/seed identity and current critical-evidence semantics.

Acceptance criteria

  • Replay the pinned Dog Ink Adventure, The Intercept, and Heresy II families under matched state, depth, time, memory, and disk ceilings.
  • Dog and Heresy each produce at least one useful reopenable result window before consuming 50% of their declared time or memory envelope, or return a checked finding explaining why that is impossible.
  • The Intercept does not regress runtime/assertion/goal/authored-knot/visible-outcome evidence against the fixed-window baseline at matched states; terminal diversity remains reported separately.
  • The policy makes identical window decisions for identical source, config, seed, posture, and telemetry ledger.
  • Tests cover tiny budgets, very slow states, steep memory growth, missing telemetry, partial terminal windows, and aggregate ceiling boundaries.
  • Human progress explains the bargain in plain language; MCP output exposes the compact decision and drill-down report ID.
  • Checked comparison includes both matched-state and matched-wall-clock views and leaves live promotion gated on evidence.

Non-goals

  • Predicting total search-space size or completion time for non-exhaustive stories.
  • Promoting shadow long-tail stopping as part of this issue.
  • Using AI to choose budgets or interpret story content.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions