Repository navigation
Use six settled Turns for product and benchmark replan defaults - #6039
Conversation
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
…n-cadence6 Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
评审对象:#6039,精确 head 59cfaf4c9cdd7ee2cb82ca9dc85180f63df205f1;基线 4bdcb2eeed64582b99cfc46b213ed755839a69b8。审查期间 head 与 base 更新,已读取新的完整22文件差异,重跑实际入口并重建打包界面;旧版本的结论没有发布或继承。
动机
维护长时间任务的用户和执行新 benchmark 尝试的 Agent。 此前产品每五个、benchmark 每三个已结算工作 Turn 触发周期复核;本版将两者改为六个,显式五或旧 Todo 配置仍保留原含义。 已独立验证第六次结算后的复核与继续执行,以及打包设置页的六、显式五、拒绝七和恢复继承;当前帮助文字及一个现有本地检查仍未同步。 不新增调度器、证据缓存、权限或计数 owner,不改变活动实验,不宣称模型吞吐、得分或安装采用已验证。 当前 CLI 帮助仍声称产品默认五,现有 cadence 检查失败;本环境另有182项可选 adapter 测试无法执行。
改动思路
复用原 cadence 配置与 TS 结算/重规划规则,只集中数值校验、调整默认及现有投影,不新增平行决策源。 当前 PR 的边界是既有配置、真实结算与设置旅程;帮助/回归检查的收尾仍需补齐,研究效果另行验证。 配置沿用原 Goal 整值覆盖、设备默认和产品默认的优先级,真实工作结算及重规划确认仍由 TypeScript 裁决。Python 仅集中数值校验与配置传输,benchmark adapter 仅传入已解析的周期,没有独立累计或结算权。保持旧默认不会解决较频繁的周期打断;只改 benchmark 常量无法通过共享配置接收六;新增计数器或跨轮缓存会重复已有 owner。
未来维护检查已采用同一原因的有界合并:删除三处数值校验重复。未要求无关 TS 迁移、恢复 RFC 或付费实验。仓库经验仅采用先读验收框架、走普通用户旅程的建议,未把经验召回当质量或长期收益证明。产品周期从五变六是明确行为变化,会多等一个已结算工作 Turn;独立验收和重复进展触发器仍可提前介入,显式五可恢复原产品周期。
具体改动
当前完整差异为22文件 +116/-80:既有数值策略、三个 transport caller、CLI/editor 投影、benchmark runtime 与两组 adapter 测试、公开操作/协议文档及聚焦 replan/config 测试。没有新增 capability、provider、共享状态或协议词表。
关键代码讲解
goal_vision_policy.py:17的默认与上限均为六,normalize_effective_turn_replan_threshold统一验证严格整数1–6;completed-Todo 的上限仍为五。machine_defaults.py:32、execution_profile.py与replan_history_codec.py复用该校验,保留 v0/v1、原单位、完整覆盖值和真实结算 owner。CLI 接受六;现有设置页更新中英文范围说明。replan_semantics.ts:186的writebackProjection要求先判断证据对当前源码和验收是否适用,再补缺失、陈旧或不足的验证;JSON 的 acceptance summary、证据关联 path outcome、绑定 refresh/spend、其他 typed exit 与显式验证门槛均保留。- Harbor/SForge/EdgeBench 把新 LoopX 尝试的默认三改六,并接收显式六;未改任务、评分、反馈或活动尝试。产品默认五改六,显式五和旧 Todo 单位继续有效。文档给出恢复五/三的现有配置方法。
采用修改前规范 docs/quota-allocation.md,固定版本 4bdcb2eeed64582b99cfc46b213ed755839a69b8,逐项映射:Goal Review Cadence 的同 Agent、distinct settlement、accepted-ACK reset、无结算不计数与有证据保留开放任务已实现;Explicit completed-Todo cadence 的旧单位、五项窗口与独立触发器保留;Precedence 的整值覆盖、revision-locked preview/apply 与清除恢复继承通过实际 UI/backend 验证。数值默认/range 是本 PR 明示的契约修订,不能用修改后的文档证明改动前契约已经要求六。模型效果为独立实验验收,当前未验证。
对主干的风险
[P2] 六周期改动尚未同步一处真实 CLI 帮助与现有回归检查。 当前 configure-goal --help 一处显示默认六,--execution-replan-after-turns 仍显示 product default: 5,与真实继承值六矛盾。现有 tests/control_plane/test_todo_replan_cadence.py:57 仍断言 editor 最大值五,当前 head 的完整聚焦运行为 1 failed / 115 passed;该现有函数在不可变当前 base 通过,head 失败为 6 != 5。这是本 PR 遗留的说明/断言收尾,未观察到配置存储或权限损坏。最小修复是在原 owner 更新帮助与现有测试,同时继续证明 completed-Todo 六拒绝、显式五保留和清除后的继承六;请重跑该函数及当前 cadence 回归,避免通过删断言或只选绿色子集掩盖失败。
实际正负路径已核验:真实 CLI 第六次 settlement 后产生周期复核,重复/重排 run 不增加计数,unspent/missing receipt 不触发,缺失证据不能 ACK,接受证据关联决策后原开放 Todo 继续执行。56个 TS 测试通过,十组显式1–5 profile/config/context 完整 JSON 保持。正常重建的 packaged UI 使用隔离真实 HTTP/backend:观察继承六,保留旧 Todo 二,拒绝 completed-Todo 六和 effective-Turn 七且无持久化变化,设备六、Goal 覆盖/清除、显式五及恢复六读回通过;中英文和390px视口已走查。外围 workspace discovery 是 fixture,不能证明生产列表排序/新鲜度。
Typecheck、scoped Ruff、diff check、源码 advisory 后的完整 native semantic smoke 和正常 packaged build 通过。benchmark 两组当前 head 测试为 35 passed / 182 skipped,原因是本环境缺少 harbor/sforge,未把这些跳过或作者的217项记录当独立 provider 资格。没有启动实验、付费模型或改变活动 Goal;安装采用、模型实际遵循及吞吐/得分保持未测。
语义与 CI 对齐
复用既有单位与 typed outcome,没有新的 enum、substring 状态分类或 domain-specific 核心义务。逐条比较 emitted rule 后,观察证据来源、JSON 事实、绑定结算顺序和停止条件保留;有 current evidence 的 no_change 仍可通过,空证据及更严格 acceptance/novelty 分支仍拒绝。证据重用是执行建议,数值准入和 semantic ACK 是机器义务,二者没有混淆。当前帮助文字的默认矛盾已单列。wait_for_ci=false,没有查询、轮询或等待远端 CI;本次请求修改来自当前本地复现与帮助读回。
我的整体评价
REQUEST_CHANGES。数值配置、结算和现有设置旅程已交付可观察增量,机制规模与问题相称,三处重复校验的有界整理有价值。长程取舍是少一次周期打断、产品方向复核可能晚一个有效 Turn;其他验收门槛、原配置和恢复路径保留,不能据此推断研究效果。用户体验仍有实际帮助/状态矛盾,现有必需本地检查仍失败;补齐同一默认变化的收尾,并在具备依赖的环境补充 adapter 正负验证后再审。当前不宣称 default-off 隔离已完整资格,不继承旧版批准,也不以绿色 CI 或 receipt 数替代验收。
English verdict: REQUEST_CHANGES - head 59cfaf4. The real six-settlement and packaged settings paths pass, but CLI help still advertises five and an existing cadence test fails at this exact head (base passes; head 1 failed/115 passed). Update the existing disclosure/assertions and rerun them. 56 TS cases and quality checks pass; benchmark dependencies are absent (35 passed/182 skipped), and model utility remains untested. CI was not consulted.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
|
Fixed the review findings on exact head Validation: the complete Todo/effective-Turn/machine-default test files pass 53 tests, including legacy Todo-six rejection, explicit-five preservation, and default restoration. Existing catalog/configuration coverage passes 50 tests (overlapping). Real CLI help readback, scoped Ruff and diff checks pass. No new evaluator/model run and no merge is claimed. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
评审对象:#6039,精确 head 7c4ea7b90da7222f105fd2d938d5ad20dc0f90cc,基线 4bdcb2eeed64582b99cfc46b213ed755839a69b8。已完整读当前24文件差异;此前59cf的帮助、目录和旧断言问题在当前 head 独立复核,未继承旧结论。
动机
维护长时间任务的用户,以及创建新 benchmark 尝试的 Agent。 此前产品每五个、benchmark 每三个已结算工作 Turn 触发周期复核;现在均为六个,显式五及旧 Todo 配置继续保留原含义。 已验证真实结算后的第六次周期复核、重复结算不重复计数,以及打包设置页的六、显式五、越界拒绝和恢复继承;最终帮助与目录也已同步。 本次只调整既有复核周期和证据重用指引;不改变任务、评分、权限或活动实验,不宣称模型实际遵循、吞吐或得分已经验证。
周期复核是停下来检查方向;它不等于任务完成。减少这一个触发器的次数可以保留较长的连续工作,但产品纠偏可能多等一个、benchmark 多等三个已结算工作 Turn。因此这里接受的是明确、可恢复的周期取舍,尚不能从配置变化推断模型质量或总成本改善。
改动思路
复用既有配置和 TypeScript 结算/重规划 owner,集中重复数值校验;无需新增计数器、证据缓存或权限机制。 当前交付是产品和新 benchmark 尝试的六周期配置、真实结算及设置旅程;长期模型收益保留为独立实验验收。 Python 沿用原数值配置/传输边界,TypeScript 保留历史计数、路径判定和结算权。只改 provider 默认值不足以让共享配置接收六;新增计数器或跨轮缓存会重复原 owner。
有界维护整理已应用:三个 reader 复用严格数值校验;最终 CLI 帮助和目录示例也引用共享默认,降低下一次默认调整遗漏的成本。没有扩大为无关语言迁移或恢复 RFC。仓库经验仅采用真实续作和普通用户旅程的检查建议,没有把经验召回当结果证明。
具体改动
完整差异24文件 +130/-85,涵盖既有配置/投影、TS replan 指引、四个 benchmark/runtime 入口、公开操作/协议文档和聚焦回归。
goal_vision_policy.py:25的normalize_effective_turn_replan_threshold统一验证严格整数1–6,默认为六;bool、分数和越界值仍拒绝。completed-Todo 范围仍1–5。machine_defaults.py:86、profile/history reader 复用原默认/校验,保留 Goal 整值覆盖 → 设备 → 产品的优先级,以及旧单位。CLI、目录示例和中英文 editor 现均同步六。replan_semantics.ts:179的writebackProjection先检查观察证据对当前源码/验收是否有效,再补缺失、陈旧或不足的探测;acceptance summary、证据关联路径 JSON、绑定 refresh/spend、其他 typed exit 和显式验证门槛均保留。- Harbor/SForge/EdgeBench 的新 LoopX 尝试默认三改六;显式值、non-LoopX 路径和冻结尝试保留,不改任务、评分或反馈算法。
修改前规范 docs/quota-allocation.md@4bdcb2eeed64582b99cfc46b213ed755839a69b8 已先读,逐项核验 Goal Review Cadence 的同 Agent/distinct receipt-backed settlement/accepted ACK reset,Explicit completed-Todo cadence 的旧单位与独立触发器,以及 Goal Review Cadence / precedence and clearing 的整值优先级、清除继承和 Receipt-backed settlement progress 的证据资格。数值五/三→六是当前明示的修订,不能拿修改后的文档反证旧规范原本已要求六。
对主干的风险
没有剩余阻塞项。当前 head 已修正真实 CLI 帮助的 product default: 5、catalog 的 Turn1–5/示例5,以及既有 editor max5 断言;原测试保留,53个 cadence/effective-Turn/machine-default 测试独立通过,completed-Todo 六仍被拒绝。
相同独立观察脚本在不可变 base/head 的真实 CLI/File runtime 运行:standard 和 fine 均从第五次变为第六次结算后的 guard 触发复核;重复 spend 不增加 journal,Todo 保持 OPEN。86个邻接回归、56个 TS 历史/语义/结算案例通过。最终三个文件仅说明、示例和原断言调整,核心规则/adapter 与这些证据的源码相同,另跑53个当前 head 检查。
benchmark 初轮缺可选依赖为35 passed/182 skipped;补齐隔离 review venv 的 Harbor0.24.0/SForge1.1.0 后217个 adapter/shared-runtime 案例全部通过。该证据是实际 adapter 入口配合 fake Codex,未冒称真实付费实验或 Docker solver qualification。
打包 frontend 与真实隔离 HTTP/backend 已走通:继承六、设备/Goal显式五、清除回六,Turn七和 Todo六拒绝且不写入;观察完整视口中的目标范围、当前来源和下一步。最终重新构建并实际读回正确范围;普通旅程没有新增必填信息或确认。workspace discovery 使用合成 fixture、model executable 为 false,不能证明生产排序、新鲜度或真实模型采用;中文桌面实测,英文源码/构建检查,未宣称所有窄屏状态。
当前头的 scoped Ruff、正常 packaged build、advisory 后的 full semantic/风险选定 native premerge 自动检查已记录。benchmark_sensitive 的人工 merge gate 保留,不把自动检查通过或自评改称已获合并许可。wait_for_ci=false,未查询、轮询或等待远端 CI。
我的整体评价
APPROVE。本次完成六周期配置与现有用户入口,同一数值改动的遗漏已在当前 head 修正;机制规模与问题相称,没有新增平行 authority 或状态。长期正向证据目前限于较少周期打断、重复结算不重复计数、原开放任务可继续,以及旧配置可恢复。真实模型效果、总成本和实验收益仍需独立实测;源码评审不关闭这些验收,也不代表已合并或升级本机。
English verdict: APPROVE - head 7c4ea7b. Final CLI/catalog disclosure and the existing editor regression are corrected; 53 final cadence/config tests pass. Identical real-CLI base/head observations show five versus six settlements with replay and open-task continuity preserved; 86 adjacent, 56 TS and 217 adapter cases pass with explicit source-invalidation checks. Packaged settings/recovery were exercised against an isolated real backend. CI was not consulted; benchmark manual merge gates and model-utility evidence remain separate.
Goal And Delivered Outcome
Periodic direction reviews interrupt new LoopX benchmark runs after three settled effective work Turns; product inheritance currently uses five. Both defaults become six, with guidance to reuse applicable work evidence before requesting another probe. Explicit Goal/device settings keep their unit and count, including five. TurnEnvelope is not required.
Author Declaration And Acceptance
Written by: model_agent (OpenAI Codex). Source: maintainer-requested cadence and evidence-reuse changes, governed by the existing Goal review contract in
docs/quota-allocation.md. Target base:main. Exact proposed head:7c4ea7b90da7222f105fd2d938d5ad20dc0f90cc.Behavior And Owning Boundary
The existing
todo_replan_cadenceconfiguration and TypeScript replan owners retain decision authority. Python normalizes the existing numeric option and transports configuration; a bounded refactor removes three duplicated validators. No new capability, provider, counter or protocol vocabulary is introduced.Intentional defaults: new LoopX benchmark runs move from three to six effective Turns; product inheritance moves from five to six. Effective-Turn configuration accepts 1–6. Official/non-LoopX profiles, explicit completed-Todo settings (1–5), persisted overrides and frozen running attempts retain their behavior. Clearing both Goal and device overrides restores six. An explicit five-Turn setting retains the prior product cadence.
Idle wakes, tool calls and planning checkpoints are not effective work Turns. Accepted same-Agent replan ACK already resets periodic and repeated-progress windows; acceptance gaps and other typed triggers remain independent. Their authority is unchanged.
The agent-consumed guidance requires checking current source and acceptance applicability, reusing sound evidence and probing missing/stale/insufficient evidence. It preserves evidence-linked acceptance/path outcomes, bound refresh/spend, explicit validation gates and other exit obligations. It introduces neither a cross-Turn cache nor a planning timeout.
Validation And Limits
Review corrections: catalog ranges and preview/apply commands now use six; CLI help derives its default from the existing numeric owner. Preserved the existing editor regression and updated its intended maximum. The complete three-file cadence/configuration run passes 53 cases, including Todo-six rejection and explicit-five/default-six behavior; 50 existing catalog/configuration cases also pass (overlapping coverage).
Final default change: 35 focused real CLI/configuration cases covered successfully (33 passed initially; two stale test expectations still assumed the former default, were corrected, and both reran successfully). Explicit five remains covered at Goal and device levels.
Full risk premerge: five direct checks and all 19 selected checks passed, including semantic vocabulary, CLI output budgets and public-boundary scans. The gate remains
manual_review_requiredfor benchmark-sensitive behavior; no self-merge is claimed.Scoped Ruff, changed-source advisory, diff checks and normal packaged Chat build passed.
Packaged frontend with isolated real backend: observed inherited six, previewed/applied explicit five with successful readback, then removed it and read back inherited six. Earlier range validation also rejected seven and verified six/apply/rollback. No active Goal was changed.
Prior same-PR layers: 62 replan/history/successor regressions, 42 TypeScript cases, 217 benchmark adapter cases and 75 latest-main focused CLI/replan cases passed before the final product-default correction. These are overlapping validation layers, not measured performance gains. Earlier packet growth was repaired by preserving obligations with shorter guidance; the full output-budget check now passes without a raised ceiling.
Full installed benchmark-worker adoption, live model behavior and sustained experiment acceptance remain separate admission checks. Failed, passed and untested layers are distinguished above.
Product Journey And Delivery
Frontend: existing capability select/count/preview/apply/remove controls, with English/Chinese range descriptions; no layout or public first-screen change. CLI: configure help and inherited default. Lark: no authority or entrypoint change. The packaged settings journey exercises the real shared backend and preserves invalid-input recovery.
Direction: existing roadmap S11 benchmark evidence and replan behavior. Future-facing pass: applied the bounded shared numeric validator refactor; no separate counter owner was added. All commits are signed off. Private state, raw trajectories, credentials, local paths and generated logs are excluded. Core changes remain for exact-head review and maintainer merge; integration experimentation is separate.