@@ -549,6 +549,56 @@ environment digests, parity status, disposition, reason codes, and redacted
549549evidence references. A public receipt cannot upgrade an unknown private audit
550550to ` eligible ` .
551551
552+ #### Declared feedback and evaluation timing
553+
554+ Qualification must compare observed access with the preregistered experiment
555+ protocol. Two independent choices belong in that protocol: whether the solver
556+ may request and receive native evaluation feedback, and whether the independent
557+ evaluator runs only after solving or also samples artifacts during solving.
558+ Background sampling does not itself authorize returning scores, diagnostics,
559+ logs, or evaluator artifacts to the solver. Local solver-authored validation
560+ remains separate from access to the independent evaluator.
561+
562+ The current ` benchmark_integrity_policy_v0 ` implementation varies network access
563+ but unconditionally requires ` official_feedback_blinded ` and
564+ ` verifier_started_after_agent ` . It cannot qualify every protocol above. This is
565+ an open toolkit/runner contract gap, not evidence that permitted native feedback
566+ is cheating. Keep affected receipts unqualified until the contract and evidence
567+ are delivered; do not set either boolean to true for a run where it is false.
568+
569+ The next bounded implementation belongs to the existing ` benchmark-toolkit `
570+ owner, with shared policy decisions in its typed TypeScript boundary and
571+ provider-specific observations in the native runner. Extend the existing policy
572+ and receipt rather than adding a competing eligibility calculation in a runner.
573+ Preserve the current blinded, post-solve default and its negative cases. An
574+ explicit protocol must be pinned before admission and bound to the runner's
575+ observations; changing it after seeing outcomes cannot qualify the old run.
576+
577+ Acceptance requires all four feedback/timing combinations through the real
578+ runner and qualification entrypoint, including these counterexamples:
579+
580+ - Allowed feedback carries only the benchmark-declared response. Hidden tests,
581+ reference answers and evaluator implementation remain inaccessible. A score
582+ response cannot grant access to their backing files or unrelated trials.
583+ - Blind background evaluation uses a controller-owned artifact snapshot and
584+ evaluator. Score stores, submission endpoints, credentials, logs and network
585+ routes must not provide a feedback path to the solver, including after resume.
586+ - Runner evidence binds the actual identity, mounts, environment and network
587+ rules to the run. Missing or contradictory evidence stays unqualified;
588+ neither a clean command scan nor a declared mode proves containment.
589+ - Provider credential exclusion is an independent boundary. A credential file
590+ owned by the solver's OS user remains shell-readable even with mode ` 0600 ` .
591+ Do not attest exclusion on that basis or waive it to admit a feedback mode.
592+ - Feedback availability is the only changed factor in a feedback ablation:
593+ retain task requirements, iteration guidance, artifact selection, budgets,
594+ evaluation cadence and source pins. Record authorized harness self-repair
595+ separately from observed evaluator access.
596+
597+ This checkpoint defines acceptance, not installed support or new access
598+ permission. Runner isolation and protocol support both remain prerequisites for
599+ countable results. Diagnostic observations retain their original qualification
600+ status; future integration must not rewrite historical receipts or active runs.
601+
552602## 7. Architecture and Ownership
553603
554604``` mermaid
0 commit comments