Skip to content

Use six settled Turns for product and benchmark replan defaults - #6039

Merged
huangruiteng merged 8 commits into
mainfrom
codex/benchmark-replan-cadence6
Oct 9, 2026
Merged

huangruiteng merged 8 commits into
mainfrom
codex/benchmark-replan-cadence6

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

Periodic direction reviews interrupt new LoopX benchmark runs after three settled effective work Turns; product inheritance currently uses five. Both defaults become six, with guidance to reuse applicable work evidence before requesting another probe. Explicit Goal/device settings keep their unit and count, including five. TurnEnvelope is not required.

Author Declaration And Acceptance

Written by: model_agent (OpenAI Codex). Source: maintainer-requested cadence and evidence-reuse changes, governed by the existing Goal review contract in docs/quota-allocation.md. Target base: main. Exact proposed head: 7c4ea7b90da7222f105fd2d938d5ad20dc0f90cc.

Criterion Disposition Evidence
Product and benchmark default to six settled effective work Turns implemented real CLI settlement, capability catalog, inheritance and override tests
Preserve explicit settings, completed-Todo range 1–5 and accepted-ACK reset implemented device/Goal precedence, retry, missing-receipt and evidence-linked ACK coverage
Reuse applicable evidence without dropping acceptance obligations implemented typed replan characterization and negative validation
Improve task throughput and score untested requires separately qualified experiments; no campaign launched by this PR

Behavior And Owning Boundary

The existing todo_replan_cadence configuration and TypeScript replan owners retain decision authority. Python normalizes the existing numeric option and transports configuration; a bounded refactor removes three duplicated validators. No new capability, provider, counter or protocol vocabulary is introduced.

Intentional defaults: new LoopX benchmark runs move from three to six effective Turns; product inheritance moves from five to six. Effective-Turn configuration accepts 1–6. Official/non-LoopX profiles, explicit completed-Todo settings (1–5), persisted overrides and frozen running attempts retain their behavior. Clearing both Goal and device overrides restores six. An explicit five-Turn setting retains the prior product cadence.

Idle wakes, tool calls and planning checkpoints are not effective work Turns. Accepted same-Agent replan ACK already resets periodic and repeated-progress windows; acceptance gaps and other typed triggers remain independent. Their authority is unchanged.

The agent-consumed guidance requires checking current source and acceptance applicability, reusing sound evidence and probing missing/stale/insufficient evidence. It preserves evidence-linked acceptance/path outcomes, bound refresh/spend, explicit validation gates and other exit obligations. It introduces neither a cross-Turn cache nor a planning timeout.

Validation And Limits

  • Review corrections: catalog ranges and preview/apply commands now use six; CLI help derives its default from the existing numeric owner. Preserved the existing editor regression and updated its intended maximum. The complete three-file cadence/configuration run passes 53 cases, including Todo-six rejection and explicit-five/default-six behavior; 50 existing catalog/configuration cases also pass (overlapping coverage).

  • Final default change: 35 focused real CLI/configuration cases covered successfully (33 passed initially; two stale test expectations still assumed the former default, were corrected, and both reran successfully). Explicit five remains covered at Goal and device levels.

  • Full risk premerge: five direct checks and all 19 selected checks passed, including semantic vocabulary, CLI output budgets and public-boundary scans. The gate remains manual_review_required for benchmark-sensitive behavior; no self-merge is claimed.

  • Scoped Ruff, changed-source advisory, diff checks and normal packaged Chat build passed.

  • Packaged frontend with isolated real backend: observed inherited six, previewed/applied explicit five with successful readback, then removed it and read back inherited six. Earlier range validation also rejected seven and verified six/apply/rollback. No active Goal was changed.

  • Prior same-PR layers: 62 replan/history/successor regressions, 42 TypeScript cases, 217 benchmark adapter cases and 75 latest-main focused CLI/replan cases passed before the final product-default correction. These are overlapping validation layers, not measured performance gains. Earlier packet growth was repaired by preserving obligations with shorter guidance; the full output-budget check now passes without a raised ceiling.

Full installed benchmark-worker adoption, live model behavior and sustained experiment acceptance remain separate admission checks. Failed, passed and untested layers are distinguished above.

Product Journey And Delivery

Frontend: existing capability select/count/preview/apply/remove controls, with English/Chinese range descriptions; no layout or public first-screen change. CLI: configure help and inherited default. Lark: no authority or entrypoint change. The packaged settings journey exercises the real shared backend and preserves invalid-input recovery.

Direction: existing roadmap S11 benchmark evidence and replan behavior. Future-facing pass: applied the bounded shared numeric validator refactor; no separate counter owner was added. All commits are signed off. Private state, raw trajectories, credentials, local paths and generated logs are excluded. Core changes remain for exact-head review and maintainer merge; integration experimentation is separate.

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
…n-cadence6

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
@loopx-agent loopx-agent changed the title Use six-turn benchmark replanning with applicable evidence reuse Use six settled Turns for product and benchmark replan defaults Oct 9, 2026

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.

Request changes conclusion (author-owned PR; GitHub blocks formal self-review)

评审对象:#6039,精确 head 59cfaf4c9cdd7ee2cb82ca9dc85180f63df205f1;基线 4bdcb2eeed64582b99cfc46b213ed755839a69b8。审查期间 head 与 base 更新,已读取新的完整22文件差异,重跑实际入口并重建打包界面;旧版本的结论没有发布或继承。

动机

维护长时间任务的用户和执行新 benchmark 尝试的 Agent。 此前产品每五个、benchmark 每三个已结算工作 Turn 触发周期复核;本版将两者改为六个,显式五或旧 Todo 配置仍保留原含义。 已独立验证第六次结算后的复核与继续执行,以及打包设置页的六、显式五、拒绝七和恢复继承;当前帮助文字及一个现有本地检查仍未同步。 不新增调度器、证据缓存、权限或计数 owner,不改变活动实验,不宣称模型吞吐、得分或安装采用已验证。 当前 CLI 帮助仍声称产品默认五,现有 cadence 检查失败;本环境另有182项可选 adapter 测试无法执行。

改动思路

复用原 cadence 配置与 TS 结算/重规划规则,只集中数值校验、调整默认及现有投影,不新增平行决策源。 当前 PR 的边界是既有配置、真实结算与设置旅程;帮助/回归检查的收尾仍需补齐,研究效果另行验证。 配置沿用原 Goal 整值覆盖、设备默认和产品默认的优先级,真实工作结算及重规划确认仍由 TypeScript 裁决。Python 仅集中数值校验与配置传输,benchmark adapter 仅传入已解析的周期,没有独立累计或结算权。保持旧默认不会解决较频繁的周期打断;只改 benchmark 常量无法通过共享配置接收六;新增计数器或跨轮缓存会重复已有 owner。

未来维护检查已采用同一原因的有界合并:删除三处数值校验重复。未要求无关 TS 迁移、恢复 RFC 或付费实验。仓库经验仅采用先读验收框架、走普通用户旅程的建议,未把经验召回当质量或长期收益证明。产品周期从五变六是明确行为变化,会多等一个已结算工作 Turn;独立验收和重复进展触发器仍可提前介入,显式五可恢复原产品周期。

具体改动

当前完整差异为22文件 +116/-80:既有数值策略、三个 transport caller、CLI/editor 投影、benchmark runtime 与两组 adapter 测试、公开操作/协议文档及聚焦 replan/config 测试。没有新增 capability、provider、共享状态或协议词表。

关键代码讲解

  • goal_vision_policy.py:17 的默认与上限均为六,normalize_effective_turn_replan_threshold 统一验证严格整数1–6;completed-Todo 的上限仍为五。
  • machine_defaults.py:32、execution_profile.py 与 replan_history_codec.py 复用该校验,保留 v0/v1、原单位、完整覆盖值和真实结算 owner。CLI 接受六;现有设置页更新中英文范围说明。
  • replan_semantics.ts:186 的 writebackProjection 要求先判断证据对当前源码和验收是否适用,再补缺失、陈旧或不足的验证;JSON 的 acceptance summary、证据关联 path outcome、绑定 refresh/spend、其他 typed exit 与显式验证门槛均保留。
  • Harbor/SForge/EdgeBench 把新 LoopX 尝试的默认三改六,并接收显式六;未改任务、评分、反馈或活动尝试。产品默认五改六,显式五和旧 Todo 单位继续有效。文档给出恢复五/三的现有配置方法。

采用修改前规范 docs/quota-allocation.md,固定版本 4bdcb2eeed64582b99cfc46b213ed755839a69b8,逐项映射:Goal Review Cadence 的同 Agent、distinct settlement、accepted-ACK reset、无结算不计数与有证据保留开放任务已实现;Explicit completed-Todo cadence 的旧单位、五项窗口与独立触发器保留;Precedence 的整值覆盖、revision-locked preview/apply 与清除恢复继承通过实际 UI/backend 验证。数值默认/range 是本 PR 明示的契约修订,不能用修改后的文档证明改动前契约已经要求六。模型效果为独立实验验收,当前未验证。

对主干的风险

[P2] 六周期改动尚未同步一处真实 CLI 帮助与现有回归检查。 当前 configure-goal --help 一处显示默认六,--execution-replan-after-turns 仍显示 product default: 5,与真实继承值六矛盾。现有 tests/control_plane/test_todo_replan_cadence.py:57 仍断言 editor 最大值五,当前 head 的完整聚焦运行为 1 failed / 115 passed;该现有函数在不可变当前 base 通过,head 失败为 6 != 5。这是本 PR 遗留的说明/断言收尾,未观察到配置存储或权限损坏。最小修复是在原 owner 更新帮助与现有测试,同时继续证明 completed-Todo 六拒绝、显式五保留和清除后的继承六;请重跑该函数及当前 cadence 回归,避免通过删断言或只选绿色子集掩盖失败。

实际正负路径已核验:真实 CLI 第六次 settlement 后产生周期复核,重复/重排 run 不增加计数,unspent/missing receipt 不触发,缺失证据不能 ACK,接受证据关联决策后原开放 Todo 继续执行。56个 TS 测试通过,十组显式1–5 profile/config/context 完整 JSON 保持。正常重建的 packaged UI 使用隔离真实 HTTP/backend:观察继承六,保留旧 Todo 二,拒绝 completed-Todo 六和 effective-Turn 七且无持久化变化,设备六、Goal 覆盖/清除、显式五及恢复六读回通过;中英文和390px视口已走查。外围 workspace discovery 是 fixture,不能证明生产列表排序/新鲜度。

Typecheck、scoped Ruff、diff check、源码 advisory 后的完整 native semantic smoke 和正常 packaged build 通过。benchmark 两组当前 head 测试为 35 passed / 182 skipped,原因是本环境缺少 harbor/sforge,未把这些跳过或作者的217项记录当独立 provider 资格。没有启动实验、付费模型或改变活动 Goal;安装采用、模型实际遵循及吞吐/得分保持未测。

语义与 CI 对齐

复用既有单位与 typed outcome,没有新的 enum、substring 状态分类或 domain-specific 核心义务。逐条比较 emitted rule 后,观察证据来源、JSON 事实、绑定结算顺序和停止条件保留;有 current evidence 的 no_change 仍可通过,空证据及更严格 acceptance/novelty 分支仍拒绝。证据重用是执行建议,数值准入和 semantic ACK 是机器义务,二者没有混淆。当前帮助文字的默认矛盾已单列。wait_for_ci=false,没有查询、轮询或等待远端 CI;本次请求修改来自当前本地复现与帮助读回。

我的整体评价

REQUEST_CHANGES。数值配置、结算和现有设置旅程已交付可观察增量,机制规模与问题相称,三处重复校验的有界整理有价值。长程取舍是少一次周期打断、产品方向复核可能晚一个有效 Turn;其他验收门槛、原配置和恢复路径保留,不能据此推断研究效果。用户体验仍有实际帮助/状态矛盾,现有必需本地检查仍失败;补齐同一默认变化的收尾,并在具备依赖的环境补充 adapter 正负验证后再审。当前不宣称 default-off 隔离已完整资格,不继承旧版批准,也不以绿色 CI 或 receipt 数替代验收。

English verdict: REQUEST_CHANGES - head 59cfaf4. The real six-settlement and packaged settings paths pass, but CLI help still advertises five and an existing cadence test fails at this exact head (base passes; head 1 failed/115 passed). Update the existing disclosure/assertions and rerun them. 56 TS cases and quality checks pass; benchmark dependencies are absent (35 passed/182 skipped), and model utility remains untested. CI was not consulted.

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
@loopx-agent

Copy link
Copy Markdown
Collaborator Author

Fixed the review findings on exact head 7c4ea7b90da7222f105fd2d938d5ad20dc0f90cc: CLI help now reads the shared default, and the existing editor assertion remains with the intended six-Turn maximum. Catalog text distinguishes effective Turns 1–6 from completed Todos 1–5; preview/apply examples reuse the existing default.

Validation: the complete Todo/effective-Turn/machine-default test files pass 53 tests, including legacy Todo-six rejection, explicit-five preservation, and default restoration. Existing catalog/configuration coverage passes 50 tests (overlapping). Real CLI help readback, scoped Ruff and diff checks pass. No new evaluator/model run and no merge is claimed.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

评审对象:#6039,精确 head 7c4ea7b90da7222f105fd2d938d5ad20dc0f90cc,基线 4bdcb2eeed64582b99cfc46b213ed755839a69b8。已完整读当前24文件差异;此前59cf的帮助、目录和旧断言问题在当前 head 独立复核,未继承旧结论。

动机

维护长时间任务的用户,以及创建新 benchmark 尝试的 Agent。 此前产品每五个、benchmark 每三个已结算工作 Turn 触发周期复核;现在均为六个,显式五及旧 Todo 配置继续保留原含义。 已验证真实结算后的第六次周期复核、重复结算不重复计数,以及打包设置页的六、显式五、越界拒绝和恢复继承;最终帮助与目录也已同步。 本次只调整既有复核周期和证据重用指引;不改变任务、评分、权限或活动实验,不宣称模型实际遵循、吞吐或得分已经验证。

周期复核是停下来检查方向;它不等于任务完成。减少这一个触发器的次数可以保留较长的连续工作,但产品纠偏可能多等一个、benchmark 多等三个已结算工作 Turn。因此这里接受的是明确、可恢复的周期取舍,尚不能从配置变化推断模型质量或总成本改善。

改动思路

复用既有配置和 TypeScript 结算/重规划 owner,集中重复数值校验;无需新增计数器、证据缓存或权限机制。 当前交付是产品和新 benchmark 尝试的六周期配置、真实结算及设置旅程;长期模型收益保留为独立实验验收。 Python 沿用原数值配置/传输边界,TypeScript 保留历史计数、路径判定和结算权。只改 provider 默认值不足以让共享配置接收六;新增计数器或跨轮缓存会重复原 owner。

有界维护整理已应用:三个 reader 复用严格数值校验;最终 CLI 帮助和目录示例也引用共享默认,降低下一次默认调整遗漏的成本。没有扩大为无关语言迁移或恢复 RFC。仓库经验仅采用真实续作和普通用户旅程的检查建议,没有把经验召回当结果证明。

具体改动

完整差异24文件 +130/-85,涵盖既有配置/投影、TS replan 指引、四个 benchmark/runtime 入口、公开操作/协议文档和聚焦回归。

  • goal_vision_policy.py:25 的 normalize_effective_turn_replan_threshold 统一验证严格整数1–6,默认为六;bool、分数和越界值仍拒绝。completed-Todo 范围仍1–5。
  • machine_defaults.py:86、profile/history reader 复用原默认/校验,保留 Goal 整值覆盖 → 设备 → 产品的优先级,以及旧单位。CLI、目录示例和中英文 editor 现均同步六。
  • replan_semantics.ts:179 的 writebackProjection 先检查观察证据对当前源码/验收是否有效,再补缺失、陈旧或不足的探测;acceptance summary、证据关联路径 JSON、绑定 refresh/spend、其他 typed exit 和显式验证门槛均保留。
  • Harbor/SForge/EdgeBench 的新 LoopX 尝试默认三改六;显式值、non-LoopX 路径和冻结尝试保留,不改任务、评分或反馈算法。

修改前规范 docs/quota-allocation.md@4bdcb2eeed64582b99cfc46b213ed755839a69b8 已先读,逐项核验 Goal Review Cadence 的同 Agent/distinct receipt-backed settlement/accepted ACK reset,Explicit completed-Todo cadence 的旧单位与独立触发器,以及 Goal Review Cadence / precedence and clearing 的整值优先级、清除继承和 Receipt-backed settlement progress 的证据资格。数值五/三→六是当前明示的修订,不能拿修改后的文档反证旧规范原本已要求六。

对主干的风险

没有剩余阻塞项。当前 head 已修正真实 CLI 帮助的 product default: 5、catalog 的 Turn1–5/示例5,以及既有 editor max5 断言;原测试保留,53个 cadence/effective-Turn/machine-default 测试独立通过,completed-Todo 六仍被拒绝。

相同独立观察脚本在不可变 base/head 的真实 CLI/File runtime 运行:standard 和 fine 均从第五次变为第六次结算后的 guard 触发复核;重复 spend 不增加 journal,Todo 保持 OPEN。86个邻接回归、56个 TS 历史/语义/结算案例通过。最终三个文件仅说明、示例和原断言调整,核心规则/adapter 与这些证据的源码相同,另跑53个当前 head 检查。

benchmark 初轮缺可选依赖为35 passed/182 skipped;补齐隔离 review venv 的 Harbor0.24.0/SForge1.1.0 后217个 adapter/shared-runtime 案例全部通过。该证据是实际 adapter 入口配合 fake Codex,未冒称真实付费实验或 Docker solver qualification。

打包 frontend 与真实隔离 HTTP/backend 已走通:继承六、设备/Goal显式五、清除回六,Turn七和 Todo六拒绝且不写入;观察完整视口中的目标范围、当前来源和下一步。最终重新构建并实际读回正确范围;普通旅程没有新增必填信息或确认。workspace discovery 使用合成 fixture、model executable 为 false,不能证明生产排序、新鲜度或真实模型采用;中文桌面实测,英文源码/构建检查,未宣称所有窄屏状态。

当前头的 scoped Ruff、正常 packaged build、advisory 后的 full semantic/风险选定 native premerge 自动检查已记录。benchmark_sensitive 的人工 merge gate 保留,不把自动检查通过或自评改称已获合并许可。wait_for_ci=false,未查询、轮询或等待远端 CI。

我的整体评价

APPROVE。本次完成六周期配置与现有用户入口,同一数值改动的遗漏已在当前 head 修正;机制规模与问题相称,没有新增平行 authority 或状态。长期正向证据目前限于较少周期打断、重复结算不重复计数、原开放任务可继续,以及旧配置可恢复。真实模型效果、总成本和实验收益仍需独立实测;源码评审不关闭这些验收,也不代表已合并或升级本机。

English verdict: APPROVE - head 7c4ea7b. Final CLI/catalog disclosure and the existing editor regression are corrected; 53 final cadence/config tests pass. Identical real-CLI base/head observations show five versus six settlements with replay and open-task continuity preserved; 86 adjacent, 56 TS and 217 adapter cases pass with explicit source-invalidation checks. Packaged settings/recovery were exercised against an isolated real backend. CI was not consulted; benchmark manual merge gates and model-utility evidence remain separate.

@huangruiteng
huangruiteng merged commit c85c57b into main Oct 9, 2026
7 of 8 checks passed
@huangruiteng
huangruiteng deleted the codex/benchmark-replan-cadence6 branch October 9, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants