Skip to content

Commit 796cb29

Browse files
authored
docs(benchmark): distinguish feedback policy from evaluator isolation (#5635)
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
1 parent f512891 commit 796cb29

2 files changed

Lines changed: 84 additions & 0 deletions

File tree

‎docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md‎

Lines changed: 50 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -549,6 +549,56 @@ environment digests, parity status, disposition, reason codes, and redacted
549549
evidence references. A public receipt cannot upgrade an unknown private audit
550550
to `eligible`.
551551

552+
#### Declared feedback and evaluation timing
553+
554+
Qualification must compare observed access with the preregistered experiment
555+
protocol. Two independent choices belong in that protocol: whether the solver
556+
may request and receive native evaluation feedback, and whether the independent
557+
evaluator runs only after solving or also samples artifacts during solving.
558+
Background sampling does not itself authorize returning scores, diagnostics,
559+
logs, or evaluator artifacts to the solver. Local solver-authored validation
560+
remains separate from access to the independent evaluator.
561+
562+
The current `benchmark_integrity_policy_v0` implementation varies network access
563+
but unconditionally requires `official_feedback_blinded` and
564+
`verifier_started_after_agent`. It cannot qualify every protocol above. This is
565+
an open toolkit/runner contract gap, not evidence that permitted native feedback
566+
is cheating. Keep affected receipts unqualified until the contract and evidence
567+
are delivered; do not set either boolean to true for a run where it is false.
568+
569+
The next bounded implementation belongs to the existing `benchmark-toolkit`
570+
owner, with shared policy decisions in its typed TypeScript boundary and
571+
provider-specific observations in the native runner. Extend the existing policy
572+
and receipt rather than adding a competing eligibility calculation in a runner.
573+
Preserve the current blinded, post-solve default and its negative cases. An
574+
explicit protocol must be pinned before admission and bound to the runner's
575+
observations; changing it after seeing outcomes cannot qualify the old run.
576+
577+
Acceptance requires all four feedback/timing combinations through the real
578+
runner and qualification entrypoint, including these counterexamples:
579+
580+
- Allowed feedback carries only the benchmark-declared response. Hidden tests,
581+
reference answers and evaluator implementation remain inaccessible. A score
582+
response cannot grant access to their backing files or unrelated trials.
583+
- Blind background evaluation uses a controller-owned artifact snapshot and
584+
evaluator. Score stores, submission endpoints, credentials, logs and network
585+
routes must not provide a feedback path to the solver, including after resume.
586+
- Runner evidence binds the actual identity, mounts, environment and network
587+
rules to the run. Missing or contradictory evidence stays unqualified;
588+
neither a clean command scan nor a declared mode proves containment.
589+
- Provider credential exclusion is an independent boundary. A credential file
590+
owned by the solver's OS user remains shell-readable even with mode `0600`.
591+
Do not attest exclusion on that basis or waive it to admit a feedback mode.
592+
- Feedback availability is the only changed factor in a feedback ablation:
593+
retain task requirements, iteration guidance, artifact selection, budgets,
594+
evaluation cadence and source pins. Record authorized harness self-repair
595+
separately from observed evaluator access.
596+
597+
This checkpoint defines acceptance, not installed support or new access
598+
permission. Runner isolation and protocol support both remain prerequisites for
599+
countable results. Diagnostic observations retain their original qualification
600+
status; future integration must not rewrite historical receipts or active runs.
601+
552602
## 7. Architecture and Ownership
553603

554604
```mermaid

‎docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md‎

Lines changed: 34 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -457,6 +457,40 @@ answer 相似这一事实来推断作弊。
457457
run identity、固定 policy/environment digest、parity status、disposition、reason code 与
458458
redacted evidence reference。公开 receipt 不能把未知的 private audit 升级为 `eligible`。
459459

460+
#### 声明反馈权限与评测时机
461+
462+
完整性资格必须比较实际访问与预注册实验协议。协议包含两个独立选择:solver 能否请求并
463+
收到原生评测反馈,以及独立 evaluator 仅在求解结束后运行,还是也在求解期间采样交付物。
464+
后台采样本身不授权向 solver 返回分数、诊断、日志或评测产物。Solver 自编的本地验证与
465+
访问独立 evaluator 分开处理。
466+
467+
当前 `benchmark_integrity_policy_v0` 实现支持区分网络权限,却无条件要求
468+
`official_feedback_blinded` 与 `verifier_started_after_agent`。它不能为上述全部协议提供
469+
资格判定。这是尚未闭合的 toolkit/runner 契约缺口,不是原生允许反馈构成作弊的证据。
470+
契约及证据交付前,相关 receipt 继续保持未合格;实际为 false 的 boolean 不得填成 true。
471+
472+
下一份有界实现归既有 `benchmark-toolkit` owner:共享 policy 决策进入其类型化 TypeScript
473+
边界,provider 专属观测留在原生 runner。扩展既有 policy 和 receipt,不在 runner 再建
474+
一套 eligibility 判断。保留当前盲反馈、求解后评测的默认行为与负例。显式协议必须在准入
475+
前固定,并与 runner 的实际观测绑定;看到结果后修改协议不能让旧运行获得资格。
476+
477+
验收需要通过真实 runner 和 qualification 入口覆盖反馈/时机的全部四种组合,并包括:
478+
479+
- 允许的反馈只携带 benchmark 声明的响应。隐藏测试、参考答案及 evaluator 实现仍不可
480+
访问;分数响应不授权读取它们的底层文件或其他 trial。
481+
- 盲反馈后台评测使用 controller 持有的交付物快照与 evaluator。分数存储、提交端点、
482+
凭证、日志及网络路由均不能形成面向 solver 的反馈通路,resume 后也必须成立。
483+
- Runner 证据把实际身份、挂载、环境和网络规则绑定到本次运行。证据缺失或矛盾仍未合格;
484+
命令扫描无命中、声明了某个模式,都不能单独证明隔离。
485+
- Provider 凭证排除是独立边界。由 solver 的 OS 用户持有的凭证文件即使权限为 `0600`,
486+
shell 仍可读取。不能据此声明排除成立,也不能为了准入某种反馈模式而豁免它。
487+
- 反馈消融只改变反馈可用性:任务要求、迭代引导、交付物选择、预算、评测频率及源码固定
488+
保持一致。获准的 harness 自修复与实际 evaluator 访问分别记录。
489+
490+
此 checkpoint 定义验收,不代表已安装支持,也不新增访问权限。Runner 隔离与协议支持
491+
均仍是结果可计入的前置条件。诊断观察保留原资格状态;后续集成不得改写历史 receipt 或
492+
运行中的实验。
493+
460494
## 7. 架构与 ownership
461495

462496
```mermaid

0 commit comments

Comments
 (0)