Skip to content

fix(quota): fence selected todo content freshness - #5911

Merged
huangruiteng merged 6 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-selected-todo-freshness
Oct 8, 2026
Merged

huangruiteng merged 6 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-selected-todo-freshness

Conversation

@mikamikasuki

@mikamikasuki mikamikasuki commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

CI regression follow-up (head 7bd705112a014d8824d87386202cd2c017c69930)

Four scoped-gate successor behavior tests pass on the exact PR base 0cd547b0c442a5c0ec7f608827b02dce7aeb0aac and failed on the previous PR head ced68b0. The deferred-resume candidate lacked content_revision; exact Todo readback computed it, so the freshness guard rejected the selected context and disabled delivery. The follow-up preserves an existing digest or derives the missing digest from the candidate source text before display compaction.

Validation on 7bd7051: test_scoped_gate_successor_tool_behavior.py and test_selected_todo_capability_binding.py passed (13 passed); Ruff, Python compile, and git diff --check passed. Required GitHub checks for the updated head are running.

Goal And Delivered Outcome

  • Outcome basis / optional anchor: Self-contained concurrency defect surfaced in the current-head review of PR #5882; no separate issue is required for this routine bug fix.
  • Goal/source and gap: The selected-Todo work-context path checked the Todo ID, status and claimant after selection, but it did not bind the selected item to the full source text. A concurrent body edit between quota selection and exact Todo readback could therefore leave delivery authorized for stale instructions.
  • Observable before → after: With the same Todo ID, status and claimant, changing only its body after selection was accepted on the prior implementation. This change carries a SHA-256 revision of the full normalized Todo text from the selection projection and compares it with the exact readback; a missing or changed revision leaves work context incomplete and delivery disallowed. The regression exercises File and SQLite authority providers.
  • Issue/task and intended base: Follow-up to the separate selection-to-read freshness gap called out in PR fix(quota): avoid duplicate selected todo context #5882 review; intended base main at 8beff60.

Author Declaration

  • Written by: model_agent — OpenAI GPT-6 Luna

Implemented against

  • Specification and revision: No written source-snapshot specification; this PR addresses the independently reported race in the PR #5882 review. Intended base: main at 8beff60.
  • Criteria:
Criterion Disposition Symbol / path Test or command
Selected Todo text is bound before display truncation implemented structured_todo_item, compact_todo_group test_selected_todo_body_change_between_selection_and_readback_blocks_delivery
Exact Todo readback must match the selected text revision implemented context_readback.py, interaction_contract.ts File and SQLite concurrency regression; TypeScript mismatch/missing-revision tests
Revision remains internal and out of canonical Todo records and outward context implemented goal_todo_projection.py, context_readback.py Canonical projection and interaction-contract suites
  • Self-check before submission: Re-read the current PR template, CONTRIBUTING guidance, governance, and the fix(quota): avoid duplicate selected todo context #5882 review. Checked open issues and PRs for duplicate freshness work. Reproduced the mutation window with an isolated Goal fixture and exact Todo update; reviewed the final 11-file diff, DCO trailer, and public-safe description.

Scope And Continuation

  • Completed scope and remaining work: The quota-selected Todo path now binds the full normalized text observed during selection to the exact readback before delivery can proceed. The revision is calculated from source text before the 500-character summary truncation and is removed from the public work-context payload and canonical Todo projection.
  • Slice boundary / successor: Complete within this scope. Other writers that bypass the selected-Todo work-context path are outside this guard.

Validation

  • Tested revision: 02efe37 (base 8beff60)
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
regression_parity passed The #5882 review documented the same body-mutation window on the prior base and head. The new File/SQLite regression observes different selected/readback revisions and confirms delivery is blocked.
real_backend passed Focused Python regression: File and isolated SQLite providers, 2 passed.
unit passed Full control-plane TypeScript suite: 4,211 passed, 32 skipped, 0 failed.
static passed Control-plane TypeScript typecheck, focused Ruff, Python compile, and git diff check passed.
unit failed The full Python module in the contribution worktree had 6 passed and 2 Turn Envelope setup failures caused by missing host/scheduler/execution context. Applying the same production patch and exact test file to a clean worktree at the base passed 8/8; these two harness failures are unrelated to the selected-Todo assertions.
  • Coverage and gaps: The regression mutates the canonical Todo body after selection but before exact readback, on both File and SQLite. TypeScript tests cover changed and missing revisions. External writers that bypass this selected-work path are not covered.

Frontend / Visual Evidence

  • UI impact: none
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: synthetic
  • Attention review: N/A

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

  • Direction / acceptance reference, when applicable: Core control-plane hardening; isolated source-freshness fix.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A
  • Semantic dimensions changed, or reviewed no-impact rationale: N/A; no migration or authority schema change.
  • Provider conformance arms run: N/A
  • Read-only legacy/file/PostgreSQL three-arm rehearsal: N/A

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths.
  • I did not duplicate maintainer-owned benchmark work.
  • I kept the change scoped to the selected-Todo freshness defect.
  • I completed the visual evidence section for UI changes, or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

CI follow-up on head 39a38e8:

content_revision is selection metadata, not a canonical Todo authority field. The coordination source projection was forwarding it into the strict canonical schema, which caused the shadow/outbox failures in CI. The projection now removes that derived field at the authority boundary. Summary compaction also preserves a revision computed from full source text instead of recomputing it from truncated text.

Validation: the three focused control-plane modules report 94 passed and 2 Turn Envelope setup failures. The File-provider setup failure also reproduces on the exact PR base 8beff60; the SQLite case has the same failure class. The selected-body freshness regression, shadow/outbox capture, and Python/TypeScript bootstrap checks pass. Ruff, Python compile, and git diff --check pass.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

CI follow-up on head 54def89: the budget regression came from serializing content_revision in the default quota summary, selected-action output, and full-detail Todo records. This digest is internal selection-freshness metadata; emitting it in agent-facing output both exposed an implementation field and exceeded the CLI output budget. The follow-up removes it from quota and monitor-poll projections while retaining the revision check on the internal selected-work path. Added regressions cover default/detail JSON and monitor-poll output. Validation: the CLI budget and monitor-poll modules pass (33 tests), the dedicated CLI budget smoke passes against base 8beff60, and Ruff, Python compile, and diff checks pass.

@mikamikasuki
mikamikasuki force-pushed the codex/loopx-selected-todo-freshness branch from 54def89 to ced68b0 Compare October 7, 2026 23:14

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | xhigh

动机

依赖自动选定任务开展工作的 agent,需要确保准入后读到的完整任务要求仍是当时选定的那一版。同一 Todo 的身份、状态和负责人没变,但选定后修改了长正文最后的验收条件;基线仍准许继续,当前版本阻止普通执行路径的旧选择,重新读取并准入后恢复同一任务。独立的 legacy、File、SQLite 真实入口对照证实了这一改善,未变的长任务和同 Goal 的其他任务更新继续可用;当前 head 新增的两个本地质量检查失败仍须处理。本次不改变任务 claim、lease、quota 或发布权限,不修复绕过此选定任务路径的外部写入,也不证明模型实际采用或长时间运行收益。

改动思路

在现有任务摘要生成路径中,先对完整、按现有规则归一化空白的正文计算内容指纹,再做500字符展示截断;精确来源读取时再独立计算指纹,既有 TypeScript interaction owner 同时核验身份、状态、负责人、生命周期和指纹相等。这个指纹是文本的派生投影,不需要用户填写,也不能成为第二份任务权限或整个 store 的 revision。对整个 store revision 比较会误伤其他任务的正常更新,因此这里的文本范围合理。

The existing typed interaction owner is the proper invariant boundary; source adapters derive text fingerprints and projections remove them at strict/hot boundaries, without duplicating selection policy. Full selected Todo freshness and same-task source-failure recovery only; canonical authority, claim/lease/quota and unrelated provider lifecycles remain unchanged. Python 保留来源 IO 和确定性字段计算,决策继续在 TypeScript。协调 shadow/完整域记录边界删除派生字段,quota/monitor 热投影及返回的 work_context 也清理它。来源失败与不匹配都由原来的恢复并重新 guard 路径处理。

具体改动

精确 head:7bd705112a014d8824d87386202cd2c017c69930;真实 merge base:0cd547b0c442a5c0ec7f608827b02dce7aeb0aac。整个 PR 是14文件 +219/-6,包含8个 production 文件和6个测试,已按完整 diff 审查;PR正文所说11文件及早期02ef/39a/54def不作为当前版本证据。发布前又观察到 head 从 ced 更新为7bd,旧草稿没有发布;当前14文件的全部证据已重新核验。

接受的前置规格是 docs/reference/required-work-context.md,revision 0cd547b。C1-full-current-body:implemented,选定要求须完整且摘要不可替代长尾;真实长文本对照保留尾部。C2-fresh-readmission:implemented,正文改变须重新准入,来源失败须保留读取命令并阻止依赖交付,恢复后同一任务重新 guard 成功。C3-single-policy-owner:implemented,普通 packet 和 Turn 使用同一类型化 fulfillment owner,读取不授予 claim/lease/quota 或发布权限。

关键代码讲解

  1. summary_item.py:117 的 todo_text_content_revision 对完整 source text 按现有空白规则生成 sha256 指纹;不是对摘要、note 或 provider revision 计算,也不判定执行权限。

  2. todo_summary.py:285 的 structured_todo_item 在截断前产生可选指纹,compact/selected transport 保留它;canonical retained rows 默认关闭。这里的函数同时增长,现有维护性检查出现新的未审查项,见风险段。

  3. context_readback.py:66 的 attach_work_context 用真实来源和既有 effect runtime 完成核验,随后从返回的 selected Todo 与 work_context 中删除指纹;source 不完整时禁用 dependent delivery。

  4. interaction_contract.ts 的 projectInteractionWorkContext 复用唯一 ENVELOPED_SHA256_PATTERN,保留原 ID/status/claim/not-done/not-archive 条件,再要求内容指纹相等。缺失或不同指纹保持 unresolved read;指纹不能绕过旧条件。

  5. _compact_selected_todo 在新提交中增加缺失指纹时的 fallback,并新增22行短 deferred fixture 测试。独立 File/SQLite 的短/长 deferred 入口对照确认仍是 successor_replan_required 非交付 lane,context complete,delivery false;此 lane 不执行选定来源读取。因此这验证的是当前 replan 语义保持,不能当作 deferred 全文 freshness 已核验。

相关简化建议:resume planning 的 display codec 会丢原 content_revision,长 deferred 的 item.text 又已经被上游压缩;当前 fallback 的指纹确实不同于 canonical 完整正文。优先在现有 codec 无损携带原派生指纹,避免把有界展示文本当作完整 source。当前没有 selected-source comparison 在该非交付 lane 执行,所以我没有将这一观察夸大成恢复死锁或新增业务阻塞;它是明确的 producer 完整性边界。

其他改动分别覆盖 strict canonical shadow、goal retained summary、quota CLI、monitor-poll 与 digest matcher 单 owner 测试。实际精确 todo list --todo-id 的冷读取仍观察到派生 content_revision,不能把“热 context 已剥离”扩写为“所有 outward Todo 输出都没有该字段”;本次没有把这个附加冷元数据认定为权限或业务回归。

对主干的风险

[P2] 更新变动后的 registry I/O census。 在真实 git worktree 的同一基线/候选 workload 下,基线全树 semantic smoke 通过,head 报 context_readback.py::_source_content::codec_read:load_registry#1 的 site metadata 已改变。新增 import 移动了已有 IO site,但 tracked project_registry_io_manifest_v1.json 未同步。用现有生成器重建并审查实际变动和分类,再运行 generator --check 及 semantic smoke;不要删除检查来获取绿灯。

[P2] 对新增的函数预算失败作有证据的 owner 修复。 同一 AST 度量基线 structured_todo_item 为88条语句/31个 decision points,head 为92/33;现有上限90/60,基线 ratchet 通过,head 新增 unreviewed oversized_decision_function。这个上限是维护性回归预算,不是外部硬限制,也不是业务失败的证明。新指纹有实际消费者价值,不能为消除2条超限删除 freshness 语义。优先在同一 owner 做有刻画验证的小范围派生 metadata 归属整理;若保留增长,应按 testing-and-quality 的 budget decision guide 提供测量、消费者价值与 reviewed 决策,不盲目提高全局上限。

独立 public CLI 对照在 legacy、File、SQLite 各跑5类场景,每个版本15例:未变短/长正文、长尾在 selection 与 readback 间经原生 Todo update 改变、peer 更新、来源不可用。head 对长尾变化拒绝,基线放行;所有 fresh guard 及 Turn 回读都恢复同一任务与当前完整正文。70个 peer 的正反顺序验证展示上限和其他行不影响当前要求。注入只安排 IO 时序/故障;native store 与 TypeScript 决策没有 mock。模型实际采用及自然时间 soak 未测。新增 deferred/原生生成 resume 命令与规划 suite 已运行,结果为 19 passed in 11.93s。

48项 Python、28项 TS 和正确 root script 的 typecheck 通过。Premerge 的5个 direct、8个成功 catalog 和8个 risk smokes 通过,另2个 catalog 失败正是上述新回归。最初 typecheck 使用错误的不存在 package 路径,已改用 npm run typecheck:control-plane;最初 archive baseline 不具备 tracked-tree 扫描条件,保留失败记录并用真正 git baseline 重跑,两项检查均通过。没有把 reviewer harness 错误算给 PR,也没有查询、等待或轮询 CI。

语义与 CI 对齐

该变更复用既有内容 digest envelope 和 interaction vocabulary,派生字段不是新的 canonical authority;开发 advisory 的两个闭集合候选已人工定位消费者。当前阻塞是实际本地 census/维护性合同未对齐,不是 pending CI。修复后必须同时保留原来的完整要求、Goal 全文读取义务、失败恢复命令及 source 不授予权限的边界,再重跑相同 workload。

我的整体评价

选定正文 freshness 的功能目标在有界实路径对照中 achieved,long_horizon 与 user_experience 为 improved:阻止旧要求并能恢复同一任务,未引入额外确认或配置;这不等于长期运行或模型采纳已验收。Observable semantics 的默认加强由原 accepted contract 的 fresh admission 条款授权,未变任务和 peer 更新保持可用。整体规模/架构 proportionate、复用 owner 合理,相关未来维护性整理应限于现有大摘要函数,无需扩展框架或语言迁移。

当前 required repository checks 有两个新增、确定、已归因的失败,semantic_alignment 为 violated,所以此精确 head 是 REQUEST_CHANGES。处理这两个 owning gate 并保留独立长尾变化/长任务不变/恢复证据后可以重审;CI 成功不是本判断的前置条件。没有 merge、安装或完成父 Goal 的结论。

English verdict: REQUEST_CHANGES - Freshness fence works across legacy/File/SQLite with same-task recovery; exact head introduces stale registry-I/O census and an unreviewed constructor budget regression. 48 Python and 28 TS tests/typecheck pass; independent base/head checks attribute the two local gate failures.
Reviewed head: 7bd7051

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Review follow-up — head 2e8b94d159f31e7d0f902647f464e811b63c5609

  • Updated the project-registry I/O manifest for the existing read site moved by this change.
  • Kept the selected-text fingerprint and display normalization together in a focused projection helper, preserving the full-text freshness check while bringing structured_todo_item back within the existing maintainability ratchet.
  • Kept this regression focused on selected-Todo freshness; the previous fixture also configured an unrelated unverified Goal acceptance contract, which correctly blocked its later envelope guard. The test now verifies that a fresh guard reads the updated full Todo text.

Validation on this head: the selected-work requirements, registry-I/O census, and maintainability-ratchet modules pass (29 tests); Ruff passes; the I/O manifest check reports 290 current sites; the repository pre-merge canary passes all 18 selected checks.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Review follow-up — deferred Todo revisions

The review identified a remaining path where a deferred Todo’s full-text revision was lost before selection. Reproduction: with a deferred Todo whose body exceeds the 500-character display limit, _structured_resume_source_items produced a bounded text projection, and resume-planning plus selected-Todo projection then generated content_revision from that truncated text. The focused regression failed on the prior implementation because this digest differed from the source body’s digest.

Root cause: the resume-source summary did not derive the revision before display normalization, and the typed resume-planning codec dropped that field from its projected payload. The fix now derives the revision from the full source body and carries it through the deferred candidate projection to selected context. The normal readback boundary removes this internal metadata after validating the selected source. Deferred/replan policy and authority are unchanged.

Validation: the focused regression fails before the fix and passes after it; 77 related control-plane tests pass; Ruff and git diff --check pass; pre-merge canary passes all 18 selected checks.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Review follow-up on current head 40c97136a0557ecaa23078338f8a0dfccd9ef3d2:

  • Updated the project-registry I/O manifest and extracted the selected-text digest projection into a focused helper; the maintainability ratchet now passes. The selected-work requirements, registry-I/O census, and maintainability tests pass (29 tests), and the pre-merge canary passes all 18 selected checks.
  • Preserved the full-text revision through deferred resume planning before display compaction. The regression fails before the fix and passes on this head; 77 related control-plane tests pass, with Ruff, diff check, and the 18-check pre-merge canary passing.

The two requested follow-ups on reviewed head 7bd7051 are addressed in the commits after that review. Please re-review the current head.

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
@mikamikasuki
mikamikasuki force-pushed the codex/loopx-selected-todo-freshness branch from 40c9713 to 1783a0d Compare October 8, 2026 02:10
@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Review follow-up on exact head 1783a0d (base 44931b6): rebased the six-commit change onto current main and reran validation. The focused Python modules pass (78 tests), the focused TypeScript modules pass (40 tests), and the control-plane typecheck, Ruff, Python compile, and diff check pass. The follow-up fixes the previously identified registry-I/O manifest and maintainability-ratchet regressions, and preserves deferred Todo full-text revisions across display compaction. Exact-head GitHub checks are still running: 2 successful, 2 in progress, 3 queued, with merge-gate expected. Native Windows behavior has not been verified locally.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | xhigh

动机

运行普通 heartbeat、quota 和 Turn 的 Agent,在选定任务后读取全文时,需要确认要求仍与选择时一致。完整审查精确 head 1783a0d9b3d19eea0b80b5ce4cde3319d068330e,真实 merge base 44931b6d22a50b949d43354e6ea498fb6b68d231,结论 APPROVE,没有剩余阻塞发现;下述一项为非阻塞精简建议。原来只核对任务 ID、状态和领取人:任务最后一条要求在两次读取之间被修改,旧选择仍被放行;现在会停止依赖交付,重新准入后带着新全文继续执行。同输入独立验证覆盖 legacy、File 和 SQLite:短正文、1801字符长正文和带无关 User gate 的 deferred 恢复,均能识别正文竞争变化并恢复;未变正文、同行更新及展示上限之外的任务不会误阻塞。本次交付是选择至全文读回这一窗口的正文一致性;不替代后续操作前的刷新义务,也不证明模型采用、真实多 Goal 长程净成本或正式安装效果。

改动思路

保留唯一的 TypeScript 交互契约判定 owner;Python 只在显示压缩前编码正文摘要、读取真实来源及隔离投影 metadata,不增加配置开关、authority 字段或平行状态机。本次交付是选择至全文读回这一窗口的正文一致性;不替代后续操作前的刷新义务,也不证明模型采用、真实多 Goal 长程净成本或正式安装效果。具体链路是完整来源正文→压缩前 revision→既有摘要/恢复/选定投影→精确当前 Todo→typed owner 比对。同一 Todo 的状态、claim、归档检查保留,缺失或不同 revision 不能履行本次必读要求;错误保留读命令并设置 delivery_allowed=false。内部摘要在主路径完成比较后删除,避免给用户或 Agent 提供第二份要求。canonical shadow 输入剔除 metadata,authority 记录的 schema 不扩展。

这是当前 docs/reference/required-work-context.md 的“改变要求须重新准入”规则(Changed requirements require fresh admission)的实质补齐;roadmap 的 Requirement continuity 属于同一 owner。文档与此前完整公开反馈均在 diff 前读取;本轮作者 body 引用了旧 head,未把声明当当前证据。Python 的哈希是来源 codec,是否放行仍由 TS 决定;无需为了这次有界修复增加独立 capability 或把整个长期验收纳入本 PR。

具体改动

全差异17文件 +273/-39:9处生产路径、7处测试和1处 registry I/O manifest。todo_text_content_revision(summary_item.py:117)按既有 Todo 文本的空白归一化语义计算 sha256;_structured_todo_text_fields/structured_todo_item 在显示截断前生成摘要,compact_todo_group 默认携带内部 revision,而 retained authority fields 明确关闭。_structured_resume_source_items 在长正文截断前生成 revision,_planning_item 传递它;_compact_selected_todo 复用已有摘要,缺失时在它本地压缩前补算,覆盖真实 deferred producer。

_source_content/attach_work_context(context_readback.py:20/66)从当前真实 Todo 全文重新计算并交给 projectInteractionWorkContext(interaction_contract.ts:237);使用现有 ENVELOPED_SHA256_PATTERN,严格比较选定与当前 revision,仍检查 ID/status/claimed_by/archive。project_coordination_source 从 canonical request 剥离 metadata,compact_quota_should_run_cli_payload 从 lane next-action 删除;不把 metadata 当可写 authority。manifest 随真实 I/O 行号更新。

测试增加来源间正文竞争、fresh guard 恢复、deferred 完整摘要的传递及输出隔离;共享 digest consumer 注册复用现有单 owner。此前验收绑定的大型集成段被移出本用例,本轮仍检查接受契约的保留代码与现有 typed coverage;没有把失效绑定作为正文竞争的替代测试。

非阻塞 P2:重规划诊断仍暴露内部 revision。 只做 successor_replan_required、没有 selected_todo 全文 read 的真实 quota 输出,agent_scope_frontier.deferred_resume_candidates[0].content_revision 仍可见;ordinary 和 scoped-delivery 路径已剥离。建议在最终 CLI 投影同一 owner 的 frontier 分支删除它,并补一个非交付 replan 输出断言。它不授予执行权限、不丢失正文,当前复验仍是 delivery_allowed=false,因此不阻塞正文一致性修复。不要在比较完成前删除该字段,也无需新增持久状态或追踪任务。

对主干的风险

独立同一生产 CLI handler、真实 legacy/FileAuthorityStore/SQLiteAuthorityStore 与 TS effect runtime,在基线和当前 head 跑60组正文/来源对照,另12组领取人变更及新任务加入,共72组决定性对照。fixture 指纹逐对一致;仅在 selection→readback 时间点用真实 todo update 注入竞争,未替换来源 reader、存储或 TS 判定。主干短/长正文变化仍 complete=true/delivery=true;当前两者 false,fresh guard 同一 Todo 恢复完整新正文。带无关 User gate 的 deferred 短/长正文正常推进,长正文变化拒绝并恢复;这覆盖上一版丢失 deferred revision 的真实 producer/consumer,非 helper-only 证明。同行更新与72个无关任务超过展示上限不改变选择;真实 Goal 文件来源缺失拒绝,恢复文件与 fresh guard 后继续。再验证原 selection 读前转交领取人:主干和当前均拒绝旧 claim,fresh guard 不重选原任务;同容器新建 peer task 不扩大本次 body gate,也不阻塞原选择。

原60组较宽探针还观察非交付 deferred replan:它不要求 selected read,仍不授予 delivery,不能宣称这条路径已验收任务正文;原探针的“所有分支隐藏 revision”断言失败保留,并转为上面实际可复现的非阻塞建议。早期把 Goal 删除放在 canonical Todo read 之后也保留为不恰当故障注入,决定性来源丢失对照已统一移到真实 Goal read 前;未把旧失败改名为通过。

源码测试77 passed,shadow/outbox/adapter28 passed,TS28 passed;typecheck通过。semantic advisory 在全树检查前识别两处 content_revision 载体:选择投影的 derived metadata,复用现有 digest pattern,未新增共享 authority vocabulary。native premerge 5 direct +10 catalog +8 risk-profile smokes 全通过,零失败/人工hold;此前 registry census 与 structured_todo_item AST ratchet 失败均已在当前 head 独立通过,未提高预算、缩小 scan 或删除质量断言。没有查询或等待远端 CI。

没有修改 authority store 实现、SQL schema 或 PostgreSQL backend;真实 File/SQLite 被这次 read/projection 测试覆盖,不声称 PostgreSQL 运行验收。未启动真实模型/付费任务;前端/Lark仍消费同一普通交互契约,没有新配置旅程,本次 CLI/host packet 与真实来源读回已验,不据此宣称 packaged frontend、安装升级或多 Agent 业务已验收。

我的整体评价

保留唯一的 TypeScript 交互契约判定 owner;Python 只在显示压缩前编码正文摘要、读取真实来源及隔离投影 metadata,不增加配置开关、authority 字段或平行状态机。本次交付是选择至全文读回这一窗口的正文一致性;不替代后续操作前的刷新义务,也不证明模型采用、真实多 Goal 长程净成本或正式安装效果。这次有界 goal achieved,long_horizon improved、user_experience improved:减少带旧要求执行的风险,并以一次必要的新 guard 恢复;正常任务没有新增参数、确认或额外用户操作。无关来源/展示变化不造成全局 revision 误阻塞,正文保留一次完整 carrier,错误恢复归原 owner。实际模型质量、token/时间净收益尚未测,哈希与校验本身也有成本,因此不能把安全收益写成已测性能提速。

Future-facing pass 已以本地正文编码 helper 降低大型 structured_todo_item 的复杂度,保留 TS 判定与不可持久化 metadata 的边界;进一步精简应收拢最终诊断字段删除,而不是新增通用 revision 框架。当前头 APPROVE,完成 readback 后原生核验旧 review closeout;合并与正式采用保留维护者边界。

English verdict: APPROVE at 1783a0d. Independent same-fixture real-entry/store comparisons show stale short/long and scoped deferred bodies rejected, then fresh admission recovers; unchanged work, peer updates and display-cap counterfactuals continue. 77+28 Python tests,28 TS tests,typecheck and premerge checks pass. Non-blocking P2: internal revision remains in non-delivery deferred-replan diagnostics. Model adoption and installed long-horizon cost remain unverified.

@huangruiteng
huangruiteng merged commit 8950509 into loopx-project:main Oct 8, 2026
22 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants