feat(replan): deliver core Goal and dense typed decision evidence - #5536
Conversation
Signed-off-by: huangruiteng <huangrt01@163.com>
Signed-off-by: huangruiteng <huangrt01@163.com>
Signed-off-by: huangruiteng <huangrt01@163.com>
|
This pull request has merge conflicts with Choose the remote for the base repository, not an out-of-date fork. git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEADFor a same-repository clone whose Keep the DCO |
Signed-off-by: huangruiteng <huangrt01@163.com> # Conflicts: # loopx/control_plane/effect_runtime_handlers.ts
Signed-off-by: huangruiteng <huangrt01@163.com>
Signed-off-by: huangruiteng <huangrt01@163.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head: a5e7a54a239b310198d09c7e750ebfe4dc269a4e. APPROVE 的结论只针对这一版本及以下明确的增量边界。
动机
跨多轮研究、工程或评审的 Agent,需要找回较早有效结果、重复尝试及尚未覆盖的方向。 一个 Agent 先得到有效结果,随后反复观察,同伴又产生较新记录;原先近期窗口藏住旧证据,现在上下文保留不同结果并折叠重复,还能沿生成的命令找回被省略的方向。
14 个身份、6 个 Goal 形状各有 221 条合成领域记录时,旧有效结果和重复次数保持,可见行数至多 24;生成的回查命令读齐本 Agent 的 221 条,拒绝跨 Agent 引用且不写进展回执。
本增量交付证据可用性和有界投影,不证明付费模型决策收益、多日运行收益或整个 TS 迁移验收,也不扩展任务、额度或外部账户权限。 现有 typed observation 没有领域级 hypothesis/反证完整性判断;完整模型收益与长期成本继续由既有 S11/T3 验收承担。
改动思路
由既有完整 compact run index 生成 Agent-scoped typed context,保留不同结果与尝试方向,合并重复观察,限量展示并提供精确引用和遗漏项回查。完整 Goal objective 与当前 task/acceptance 分开;局部任务不能冒充最终目标。quota、PR review 和 handoff 复用同一 TS owner,Python 保留 source/codec,旧重复 evidence-log writer 和强制读回仪式退休,历史 reader 保留。
复审修复了一处真实恢复缺口:生成的遗漏读取曾按 Goal 先截断,再过滤 Agent,还受 60 条 decision lookback 限制。现在完整 decoded source 先按 Goal/Agent scope,再应用调用者 limit;普通 history 与 replan decision window 分开,Goal-wide quota 计算仍用完整 Goal source。
具体改动
关键代码讲解
loopx/control_plane/work_items/replan_context_codec.py:15的replan_history_from_status读取完整 compact index;source 缺失仍明确 snapshot 边界,不把有限切片当完整 authority。loopx/control_plane/work_items/replan_context_codec.py:31的replan_evidence_rows校验 Goal、Agent、时点与 typed observation,peer 不提供本 lane 的证据。loopx/control_plane/work_items/replan_context.ts:74的projectReplanContext按完整事实合并重复,保留结果/route 差异,至多 24 行;delivered表示 host 提供上下文,不表示模型消费或验收完成。loopx/history.py:324的collect_history接收独立scoped_agent_id,在 limit 之前过滤,runs/latest_runs/latest_status 一致;原 bounded decision lookback 与 Goal quota 输入保持独立。loopx/control_plane/effect_runtime_handlers.ts的两个 replan-context 方法继续 lazy load 同一现有 TS 模块,保留最新 main 的冷启动结构,没有恢复大批 static imports。
独立规格:docs/reference/protocols/goal-vision-replan-contract-v0.md,spec_revision e55489c7791f326abe63e7b9b2952a134d652cf7。Projection Contract 章节规定:Agent public-safe evidence、optional drill-down、不以读取 receipt 算进度,均有真实入口证据。Vision Continuation Audit 章节规定:当前权威证据优先,弱/旧/缺失证据不能当验收。S3/S6 是本次增量;S11/T3 模型 score gain、全调用者迁移及长期成本保留既有后续阶段。
对主干的风险
原反例:多 peer 下,一个 Agent 的 221 条记录被较新同伴记录挤掉;生成的成功回查只能给 0–180 条 top-level 或最多 60 条补充,多数旧方向缺失。修复后 14 个身份逐个沿 generated read_action 实际执行 CLI,全部读齐自身 221 条,普通 limit=10 正确,latest_status 不丢;所有 18 条 omitted 方向可回查。wrong-Agent exact ref 拒绝,读操作不改 index、不产进展或 read-gate receipt。
领域数据来自完整真实 authority 的 Goal/任务语境,结果观察是合成的:评审 head drift、lease/停止、产品重启、运营来源、决策未知背景、笔记索引、系统失败反例、工具冷暖启动、交接策略、发行人披露/反证、量纲与权益冲突、Bot 排队修正。14 个身份共 3094 条领域历史;旧 advanced 保留、180 次重复合并、42 个 distinct 仍在完整 novelty index,可见最多 24。上下文不会仅凭“盈利/已停止/已完成”文本授予领域结论或执行权。
修复入口聚焦 31 passed,扩展回归 179 passed,最终 premerge canary 19 项均通过;集合有重叠。TS 聚焦 44 passed、core typecheck、quota/review/handoff 实际入口与 CLI 差分通过。原完整差分为 102 baseline / 96 head、0 candidate-only;退休 6 项入口和 12 项 review-required intentional deltas 明确披露。两项 canonical periodic successor 原先在同一 immutable base 与原 head 具有相同 test id/guard 错误,已单独归因;不是 omitted-history 反例。当前 Goal 的原生评审/合并契约明确 wait_for_ci=false,已停止轮询;本次批准依据当前 head 的本地入口、反例、差分与 canary 验证。此前已观察到 DCO 和真实 PostgreSQL 检查成功,其余 GitHub CI 尚未全部完成,不宣称全绿;发布/部署和非要求版本路径的 skip 不当通过。用户已明确授权两 PR 自合并,仍须公开 exact-head review、无未解决 review thread 和原生 ready=true。
语义与 CI 对齐
默认上下文与 CLI 入口变化已披露,未宣称 opt-in/default-off。typed observation 决定结果分类,未通过 substring 或自然语言判断领域真假。evidence 提供、读取、进度与最终验收分层;旧 evidence-read 义务已退为可选诊断。snapshot 不等于全 source,未知 source 不当完成。manifest 漂移在修复后重生成并复跑,保留完整扫描根和真实分类;公开边界 clean。
成本明示:同样 crowded quota fixture 由 26844/718 增到 33523 chars/806 lines,约 25% 字符增长;相应 crowded regression guard 为 34000/830,普通/非 replan guard 保持原界。diagnose 43132/804 对应 44000/850。新增内容保留完整 Goal 与证据差异,重复观察被压缩且有可用 drill-down,增量预算不是无限豁免。每个领域 context 约 23.2–26.1k JSON 字符;本地 warm helper 约 12–15ms,仅是 helper 成本。完整索引 IO 随历史增长,未测付费模型 token/决策收益或端到端加速。
我的整体评价
有界且可回查的完整证据比单纯扩大最近窗口更适合长程工作。Future-facing pass 集中 TS 规则、退休重复 live entrypoint、保留真实持久化 reader,修复 scope-before-limit 并保持 lazy handler。预计效果正向,尤其证据稀疏或经常有 peer 并行活动的 lane;效率短期每次输入成本负向,长期减少重复探索的净收益尚待观察。不能把 context 到达等同模型采用。
English verdict: APPROVE — exact head a5e7a54 delivers diverse bounded evidence and a working Agent-scoped omitted-history recovery path, retains the current lazy runtime architecture, and preserves authority boundaries. Dense-context output cost is explicitly measured; sustained model-effect and multi-day cost claims remain outside this increment.
Recent history slices can hide an earlier useful result behind repeated observations and newer peer activity. Project the full Goal objective separately from current task acceptance, and deliver Agent-scoped typed evidence that retains result/route diversity, folds duplicates, bounds display to 24 rows, and provides exact references plus omitted-history recovery.
Quota, PR review and handoff share the TypeScript context owner over the existing compact run index; Python adapts source IO. Retire the duplicate live evidence-log interface and mandatory diagnostic-read ritual, retaining historical readers. Agent-scoped history now filters the complete decoded source before the requested limit; bounded replan decision lookback and Goal-wide quota accounting remain separate. Preserve the current lazy runtime handler architecture.
Validation: 14 authority-derived identities across six Goal shapes, 221 synthetic domain observations each. Early results, 180 repeated observations, all distinct routes, exact-reference refusal, ordinary scoped history and the generated omitted read command are exercised through the actual CLI/index backend. Every scoped omitted read returns all 221 records without index writes or progress receipts. Focused context/history tests, extended regressions, TS/typecheck, real CLI base/head differential and all 19 latest-main premerge canaries pass.
Cost tradeoff: the same crowded quota fixture grows from 26844 chars/718 lines to 33523/806 (about 25% more characters); the documented crowded regression guard is 34000/830 and ordinary/non-replan guards are retained. Full source IO still grows with history, even though visible evidence rows are bounded. S11/T3 sustained model-effect and cost qualification remains open; this increment does not claim token savings, paid-model quality gains or new execution authority.
Final exact-head GitHub checks and published review determine merge readiness.