feat(delegation): stop delegated members with acknowledged, settled receipts - #5308
Conversation
Add `lock_holder_liveness` to the file-lock owner and `turn_lane_liveness` on top of it. Both classify a lane's last executing Turn from its holder record alone: `released` on a clean exit, `dead` when the record names this machine and the pid is gone, `foreign_host` when the pid cannot be checked here, `unreadable` when a lock file carries no parseable record, and `live` only when a same-host pid is still alive. The probe never touches the kernel lock. A probe that acquired it for an instant would refuse a real `run-once --execute` racing that instant with `turn_lane_in_flight` for nothing; the new test drives the real fence wrapper concurrently with a continuous probe and proves the Turn is admitted exactly once. The holder host label is now single-sourced so the writer and the readers cannot disagree on what "this machine" is. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Delegated members had no stopped observation: prepared, running and turn_returned could only end in accepted or rejected. Add "stopped" as a terminal observation reachable from the three open states and keep the inventory check driven by the same transition table. Add the exported collaboration.delegation.stop decision: a stop request settles only with an acknowledgement from a process that held the operation lock plus a free operation lock and a free Turn lane lock. Free locks with no acknowledgement are "unknown", never a fabricated settlement, and a grace timeout alone changes nothing while a lock is still held. Register the RPC handler next to the existing delegation observation handlers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…eceipts
Delegated members could not be stopped: the worker held the operation lock
for the whole run and rewrote the execution record from memory, so nothing
written into that record could reach it or survive it, and a killed run left
its Turn and hard lease unaccounted for.
Add a stop receipt beside the execution record (executions/<h>/<op>.stop.json)
written only under the existing .dispatch lock, never into the record. Its
phases are requested -> acknowledged -> settled with the terminals unknown and
noop, and it records the requester, the lock-holding worker (pid, process
group, host), the acknowledgement (pid, observed status, Turn key), the lease
release outcome and the settlement facts (operation lock, Turn lane lock,
Turn journal status).
Delegations.stop: a terminal record returns an identical noop receipt on every
call. Otherwise the request is written; when no worker holds the operation the
requester takes the lock, marks the record stopped, releases the hard lease
and settles. A same-host holder is SIGTERMed by process group (its run-once
child and host bridge follow) and SIGKILLed only if it still holds the
operation after the grace; another host's holder is never signalled and finds
the request itself. Settlement is the typed collaboration.delegation.stop
decision over lock facts: elapsed time is never a receipt.
Worker: a SIGTERM handler raises DelegationStopRequested once; checkpoints
before running, before each run-once and before completing the Todo raise on
a stop file; the record is written only through a fenced write that re-reads
the stop file under .dispatch and raises DelegationFenced for a stop this
process did not acknowledge, so a late-returning or other-host worker records
no Turn result and completes no Todo. Acknowledgement marks the record
stopped, releases the hard lease and leaves the in_progress Turn journal for
inspection. execute records the worker's pid/pgid/host at entry and
acknowledges from under the lock when a stop already exists.
resume refuses stopped work ("start a new operation id"), wait returns on
stopped, and read exposes the stop phase. file_lock gains local_lock_host and
read_lock_holder so the stop names the holder exactly as lock records do.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Add `loopx delegation stop --operation-id ID --execute` next to start/resume/adopt (--execute is required for the same reason) and the `stop_delegation(operation_id)` MCP tool, which runs the blocking stop off the event loop like wait_delegation. Both return the same receipt and never resume or rerun work. The inventory reader skips `<op>.stop.json` sidecars, which sit beside execution records but are not records, and the delegation context and subagent context projections count `stopped` receipts instead of dropping them. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Semantics-first coverage for stopping delegated members: - stop while a detached worker runs a sleeping fixture host: the worker acknowledges SIGTERM under its own lock, the receipt settles only with a free operation lock and a lockable Turn lane, the record bytes stay frozen afterwards, the Todo stays open, the host and worker processes are gone, the Turn journal stays in_progress, and resume is refused without spawning; - stop after accepted: noop, identical on repeat, no stop file and artifacts unchanged; stop without a holder is acknowledged by the requester; - a worker SIGKILLed before acknowledging settles as unknown, not stopped, and resume stays refused; - a fenced write after another process's acknowledged stop raises and writes nothing, while an unacknowledged stop is taken from under the lock at entry; - CLI stop requires --execute and repeats its receipt; inventory pages past stop sidecars and reads a stopped record as stopped; - TS: stopped is terminal and reachable only from open observations, and the stop decision settles only on acknowledgement plus free locks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Describe `delegation stop --execute` / `stop_delegation` in both reference documents, in English and Chinese: where the request lives, how a same-host worker is signalled and acknowledges, why another host's worker is left to find the request, what settled, unknown and noop prove, and that stopped work needs a new operation id while its Turn journal and open Todo remain for the coordinator to inspect. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
… lock The stop settlement probed the member's Turn lane by acquiring its lock for an instant, which could refuse a legitimate Turn of the same member racing that instant with turn_lane_in_flight. It also carried its own host label helper next to the file-lock owner's. Settlement now reads the lane's last holder record through turn_lane_liveness: released, dead or absent frees the lane; a live same-host holder frees it only when it sits outside the recorded worker's process group; another host's holder, an unattributable holder or an unreadable record keep the stop open. The operation lock is read the same way, so a stop never refuses a legitimate resume or status read. The local host helper is deleted in favour of lock_holder_host_label. A regression test races a real lane acquisition against settlement and proves the Turn is admitted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Regenerated with scripts/generate_project_registry_io_manifest.py after rebasing onto main. Site ids and classifications are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
d5aafc8 to
5e7b7bf
Compare
huangruiteng
left a comment
There was a problem hiding this comment.
精确 head:5e7b7bfd8fb425312d25a771911984040ac9ee03。按 LoopX PR review capability 检查整份差异,未把相邻 PR 的发现移植到本 PR。
动机
Roadmap R2 要求 stop/cancel/restart 保留工作并 fence 旧执行者。已有 delegated member 无可用停止入口,杀进程也不能证明结果、Turn 和租约已妥善收尾。本 PR 增加 CLI/MCP 停止与持久 ACK,是合理的有界交付;但用户需要的是“停止已完成”的可信结果,不能在返回 settled 后原执行宿主仍可继续工作。
改动思路
复用既有 binding、GoalRef、operation 单飞锁、dispatch 锁及 Turn lane owner。stop sidecar 表达不可由普通运行状态推导的明确停止意图;worker 在自己的锁下 ACK,fenced writer 禁止较晚的结果/完成写入。TS owner 持有 requested/acknowledged/settled/unknown/noop 决策,Python 适配信号、权威租约释放和日志 IO。CLI delegation stop --execute 与 MCP 返回同一收据,read/wait/inventory 消费持久事实;它没有授予跨 host 信号、跨 peer 权限或 Goal 结算权。
具体改动
整份差异为 17 文件、+1222/−80:停止入口、typed lifecycle、operation worker、lane/holder 只读观察、上下文/清单、派生 registry 清单、中英文档与真实进程测试。检查范围包含随分支引入的 lock/lane seam,而不是只看 stop helper;没有引入相邻 PR 的 execution-facts 变更。sidecar 有明确生产者、requester 身份与锁下更新路径,停止后 resume 拒绝重新运行;未写 stop 意图的外部 SIGTERM 保留原 recoverable 路径。
关键代码讲解
- Delegations.stop 校验现有 operation/binding,锁下保存停止意图,只向可归属的同 host worker 进程组发信号,终态返回稳定 noop。
- _acknowledge_stop 在 operation 锁内标记 stopped、记录 ACK,清理 bootstrap 并尝试释放已有硬租约;ACK 本身不是宿主清理完成的证明。
- _settle_stop 只提供 operation lock free 和 worker lane released 两项事实,再调用 decideDelegationStop;当前缺少实际 native host/后代已 drain 的事实。
- native host process cleanup 已有独立进程组与异步清理 owner,SIGTERM 后有 300ms grace 再 SIGKILL。worker/lane 释放不能替它宣告清理完成。
对主干的风险
[P1,阻断] settled 可早于实际宿主和后代终止。 在 File、SQLite 两种真实权威后端,我沿用现有 native fixture,令宿主及其同组子进程忽略 SIGTERM,其他路径仍是实际 detached worker → CLI run-once → native Node bridge → host transport。调用 stop 后约 0.26s 返回 phase=settled、operation lock free 和 lane released,但用 macOS ps 独立检查,宿主与子进程此刻都仍存活。native cleanup 稍后会杀掉它们,因此这是“过早结算”,不是声称永久 orphan;这段窗口内旧效果仍可能执行,接续工作也可能与其重叠。
原因是 _signal_worker 的提前返回 与 _settle_stop 把锁释放当成 drain;native bridge/host 位于其他进程组,拥有独立的异步清理生命周期。请复用现有 native host cleanup/readback owner,只有实际 owning host/后代清理完成才能结算;缺少可归属证明时保持 acknowledged/unknown 并允许同身份读回恢复。不要用固定 sleep、扩大超时或跨 host kill 冒充证明。补一个“宿主和同组后代忽略 TERM”的真实 File/SQLite 回归,断言返回 settled 的那一刻两者已退出,并覆盖清理中断/重复读取不产生重复接纳或完成。可放入 tests/test_local_delegation.py,重跑本段所列 pytest/TS 验证。
现有测试只在 stop 返回后再等待退出;其 /proc 检查在 macOS 读不到路径时还会把存活误当退出。请一起改为平台有效的进程检查及即时后置条件。PR 正文也应同步实现:Turn lane 是只读 liveness,operation lock 在 stop 已写后仍有短暂 kernel probe,并非全部“no lock-taking”。
语义与 CI 对齐
亲自运行 66 项 Python、18 项 TS、控制面 typecheck 均通过;独立 native drain 反例在 File/SQLite 各失败一次。这证明现有绿测试没有覆盖上述后置条件。标准 canary 14 个选择检查及 5 个直接检查中,只有 twin-budget 检查失败:原始同一 full-tree budget owner 在不可变基线 902b99698050fc21af2c3a8b0d0d3b8561d0bdd7 和本 head 都报 44/43;扫描 owner、预算和 twin 列表未被此 PR 改动,故归为既有无关失败,不作为本次 request changes 的理由。基线完整 smoke 另有旧 registry 清单行号失配,未掩盖或称其通过。
七种普通 status/Goal Chat 路径在同一夹具下,完整基线/head 观察一致;new stop 是明确请求才触发,安装、帮助发现或普通 SIGTERM 不自动激活停止。远端 CI 未查询,打包 App(无 stop 控件)、Lark、跨真实 host 及真实模型未验证。复用既有 delegation vocabulary 的范围恰当,但 settled 目前违反其停止完成语义,不能由文档描述消除。
我的整体评价
REQUEST_CHANGES。长期推进和用户体验均存在具体回归风险:用户看到结束收据时旧执行仍在 drain。规模与 CLI/MCP 到持久 ACK/读回的 R2 切片大体相称;future-facing 检查建议把 stop 协调收敛到现有 collaboration 边界、复用 native cleanup 事实,避免 Python 再建一个 host 生命周期 owner。较大的模块拆分可在有 characterization 时另作有界整理,不作为本次 LOC 门槛。当前必须先修过早结算并用原生负例证明修复,再复审精确 head;保留 rollback 对 stopped 历史的 caution,不授予自合并权限。
English verdict: REQUEST_CHANGES — exact head 5e7b7bf. A real File/SQLite native-process counterexample returns settled while the owned host and descendant remain alive. The pre-existing 44/43 twin-budget failure is unrelated and is not the blocker.
…up exit The worker and its Turn lane let go while the Host supervisor is still terminating the Host, so their release never proved that the old executor stopped. The Host transport now names the process group its supervisor owns in a record beside the operation, the drain is read back from that record without signalling anything, and the typed decision requires an exited group before a stop may settle. A group that cannot be attributed keeps an acknowledged stop open for a later same-identity read, and the TS supervisor stays the only owner that terminates a Host. Signed-off-by: song <liusongstep@gmail.com>
A real File/SQLite fixture whose host and same-group child ignore SIGTERM asserts both are gone at the instant settled returns, with a platform-valid process check in place of the /proc read that treated a live macOS process as exited. Interrupted cleanup, repeated reads, an unacknowledged holder and an unrecorded group get their own negative coverage. Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
New exact head: [P1, blocking]
|
cocolord
left a comment
There was a problem hiding this comment.
动机
这个 PR 要解决的是一个真实且高价值的问题:已派发的 delegated member 没有可验证的 stop,单纯杀 worker 也不能证明 Turn、原生 Host、后代进程和租约已经安全收尾。新 head 609530d2fb64e86ae71eed13ce22d4abdfd7150e 已修复我上次指出的 native Host 提前结算问题;File/SQLite 的真实进程组测试现在能证明返回 settled 的瞬间 Host 与同组 child 都已退出。
改动思路
实现复用现有 operation 单飞锁、dispatch 锁、Turn lane holder、TS collaboration decision、native Host supervisor 和 task-lease authority。stop.json 保存不可由普通运行状态推导的停止意图,host.json 保存 supervisor 产生的进程归属投影;worker 在 checkpoint 或 fenced write 处确认,decideDelegationStop 再根据 ACK、operation/lane 释放和 Host drain 选择 requested/acknowledged/settled/unknown。CLI 与 MCP 共用同一入口,read/wait/resume/inventory/context 共用持久事实。
具体改动
整份差异为 21 个文件、+1596/-81:12 个 runtime 文件增加 stop CLI/MCP、typed stopped/phase 决策、worker 信号与写 fence、holder liveness、Host spawned 回报/原子 sidecar、租约释放及各投影;7 个测试文件覆盖 TS 转移、真实进程组、File/SQLite、CLI/inventory/lane;2 个文档文件补充中英文使用与回滚语义。
关键代码讲解
Delegations.stop写入一次 stop intent,只向可归属的同机 worker 进程组发信号,并复用同一stop_id读回。_settle_stop汇总 operation lock、Turn lane 与host_process_drain;这是本轮对旧 blocker 的有效修复。decideDelegationStop是 typed phase owner,超时本身不制造 terminal receipt。runHostProcess在 Host 接收输入前报告 owned group;报告失败会取消,TS supervisor 仍是唯一 kill owner。_execute负责最终 Todo、reply 与 accepted 写入;当前剩余 blocker 都集中在 stop 与这些既有 authority 的交界。
对主干的风险
[P1,阻断] stop 与 Todo 完成/结果发布没有共同线性化边界。 _execute 在 1557 行 只做一次 pre-check,之后 _complete_delegated_todo 与 return_result 都在 dispatch fence 外提交,最后 _observe/_fenced_write 才重新看 stop。我在 File/SQLite 让 worker 停在 Todo effect 内,先通过公开 stop 写入请求,再放行 worker;两种后端都观察到 todo_completed、reply_published,随后却返回 phase=settled,status=stopped。这违反“late/other-host worker 不完成 Todo、不发布结果”的核心承诺,也可能让 successor 与已提交效果重叠。请让 stop 与两个外部 effect 通过同一 authority fence/CAS 决定谁先提交,并分别加入竞态回归。
[P1,阻断] required hard lease 释放失败仍会 terminal settle,且不会重试。 _settle_stop 的 typed 输入没有 lease release;File/SQLite 负例中 release authority 报错后仍得到 settled 和 lease_released=false,第二次 stop 原样返回且不再尝试释放。此时同 operation 已禁止 resume,但 Todo 最长仍会被 45 分钟 hard lease 阻塞。请把 required lease release 纳入 typed settlement,保留同一 stop_id 的可重试/可操作恢复路径,并覆盖失败后恢复。
[P1,阻断] Windows 上已启动 Host 的 stop 没有收敛或 fail-fast 路径。 host_process_drain 在没有 os.killpg 时恒为 unattributable;同机 signal 也依赖 killpg。通过真实 public stop 决策执行该平台分支后,即使 Host record 已是 finished,File/SQLite 都在重复调用中永久保持 acknowledged/host_process_drain_unproven。仓库已有 Windows Host tree-best-effort 支持,文档却没有把 stop 限定为 POSIX。请补 Windows supervisor/tree 完成事实和 CI,或在写入任何 stop/status 前明确 fail fast 并文档化;不能留下不可恢复的 open receipt。
验证方面,旧 blocker 的 8 个 File/SQLite native drain/cleanup-interruption 用例、另外 8 个 stop 用例、29 个 Python host/lane/CLI/inventory 用例、29 个 Node delegation/Host 用例和控制面 typecheck 均通过。远端 shard 1/3/4 的三项失败,我用同一命令在 immutable base 3ec049e138917a8cce4f84197ba196d26445b2b0 与本 head 都复现为相同的 generated-twin / prompt-upgrade 断言,因此是独立 merge-readiness hold,不是上述 request-changes 的依据。
语义与 CI 对齐
stopped 和 settled 是新的公共终态,不是提示性文案;它们会驱动 resume 拒绝、wait 返回、inventory/context 和后续协调决策。当前 typed decision 没有外部 effect 或 required lease 的事实,Windows unattributable 又没有恢复转移,因此实际语义仍宽于实现。最小修复后请重跑 File/SQLite effect race、lease failure/retry、Windows CLI/MCP stop、现有 native process suites 及完整 required CI。
我的整体评价
REQUEST_CHANGES。这个方向和范围有明确收益,stop sidecar、typed owner、只读 lane liveness 以及新的 Host group readback 都放在合理边界;上一轮 blocker 也确实修好了。但 whole-PR 的价值取决于 settled 能否成为后续继续工作的可信边界,目前 exact head 仍允许 late Todo/reply、不可重试的 lease failure,以及 Windows 永不收敛。长期推进和用户体验都因此 not_yet_proven。建议在现有 effect/lease/Host owner 上做有界修复,不再增加平行生命周期;修复后我会按新 exact head 重跑同一组正负路径。
English verdict: REQUEST_CHANGES — exact head 609530d2fb64e86ae71eed13ce22d4abdfd7150e. The prior native-Host early-settlement bug is fixed, but stop can still settle after Todo/reply effects commit, failed required lease release is terminal and never retried, and launched Windows Hosts have no converging or fail-fast stop path. Focused submitted suites pass; the three red repository shards reproduce unchanged on base and head and are unrelated.
Signed-off-by: song <liusongstep@gmail.com>
Re-review request — exact head
|
cocolord
left a comment
There was a problem hiding this comment.
动机
这个 PR 要解决的是一个真实且高价值的问题:已经派发的 delegated member 需要一个可验证的 stop 边界,不能只杀 worker、留下 Turn、原生 Host、Todo 或 hard lease 处于含糊状态。当前 exact head 796f688674c3a404daec89cfa6d3b5801683e292 只是把最新主干合入上次已审的 head;目标收益仍然明确,但只有当 settled 能证明后续继续或重派不会与旧执行重叠时,这个收益才成立。
改动思路
实现把不可由运行记录推导的 stop intent 放进独立 stop.json,把 Host supervisor 观察到的进程归属放进 host.json;CLI/MCP 通过 Delegations.stop 发起请求,worker 在 checkpoint 或 fenced write 处 ACK,TypeScript 的 decideDelegationStop 再组合 operation lock、Turn lane 和 Host drain 事实。正常 POSIX 路径是 requested → acknowledged → settled,read/wait/resume/inventory/context 复用同一持久事实;异常路径应当在 effect、lease 或平台清理尚未闭合时保持可恢复、不可误报 terminal。
具体改动
相对 exact base fd5f31bb3ad57448c32df6c7cddde36d07446934,整份差异为 18 个文件、+1322/-53:10 个 runtime 文件增加 stop CLI/MCP、typed stopped/phase 决策、worker 信号与写 fence、Host 进程组 sidecar、lease release 读回及 inventory/context 投影;6 个测试文件覆盖 TS 转移、真实进程组、File/SQLite、CLI 和 inventory;2 个中英文参考文档说明 stop、receipt 与回滚。先前 stacked 的 lane-holder liveness 已进入 base,本 PR 当前复用它。
关键代码讲解
Delegations.stop写入一次 operation-specific intent,只信号可归属的同机 worker,并把重复调用收敛到同一 receipt。_settle_stop读取 operation lock、Turn lane、Host drain 后调用 typed decision;目前 required lease release 并不是 decision input。decideDelegationStop是 requested/acknowledged/settled/unknown 的状态 owner,正确地拒绝用超时本身制造 settlement,但没有外部 effect 或 lease-success 事实。host_process_drain在 POSIX 上只读验证 bridge/process group;没有os.killpg时恒为unattributable。_execute的最终 effect 段 依次完成 Todo、发布 reply、再通过 fenced_observe写 accepted;stop 只在 effect 前做瞬时 pre-check。
当前 head 的第一父提交就是上次审过的 609530d…,第二父提交是 exact base;18 个 PR 文件里只有 delegation.ts、delegation_context.py、effect_runtime_handlers.ts 吸收了主干变化。下面三个 blocker 的 Python owner、Host drain 和测试文件均 byte-identical,因此不能把合主干视为修复。
对主干的风险
[P1,阻断] stop 与 Todo completion / result publication 仍没有共同线性化边界。 worker 在 1557 行 做一次 pre-check,之后两个外部 effect 都在 dispatch fence 外提交,最后写 execution record 时才再次看到 stop。exact-head 的 File/SQLite 反例都先持久化 stop,再放行 worker,仍观察到 todo_completed 和 reply_published,最终 receipt 却是 phase=settled,status=stopped。请让 stop 与两个 effect 通过同一 authority fence/CAS 决定先后,并分别提交 race regression。
[P1,阻断] required hard lease 释放失败仍被 terminal settle,且不会重试。 _settle_stop 没把 lease.released 交给 typed decision;两种后端注入一次 authority failure 后都返回 settled, lease_released=false,第二次 stop 原样返回且 release attempt 仍只有一次。此时 resume 已禁止,Todo 却可能被 hard lease 阻塞到 TTL。请把 required release 纳入 settlement obligation,并让同一 stop_id 在失败后可重试或进入明确可操作的恢复态。
[P1,阻断] Windows 上已启动 Host 的 stop 仍没有终结或 fail-fast 路径。 host_process_drain 在缺少 killpg 时恒返回 unattributable;File/SQLite 的公开 stop 路径即使读取到 phase=finished 的 Host record,重复调用仍永久停在 acknowledged/host_process_drain_unproven。通用 windows-powershell 绿灯没有覆盖这个新 stop contract。请提供 Windows supervisor/tree completion 事实并在 Windows CI 覆盖公共路径,或者在写 stop/status 前明确 fail fast 并文档化平台边界。
验证上,官方 exact-head Python focused suite 63 项、Node delegation/Host 29 项、control-plane typecheck、Ruff、diff hygiene、docs-governance 与 semantic-vocabulary smoke 均通过;独立的 6 个 File/SQLite 负例也稳定复现上述三个错误结果。绿色正向覆盖说明主体实现可运行,但没有覆盖 stop 在最后 checkpoint 之后与外部 effect/lease/platform清理交错的语义。
语义与 CI 对齐
stopped 与 settled 是机器消费的公共终态,会驱动 resume 拒绝、wait 返回、inventory/context 和后续重派,不是“guidance”。当前 typed contract 未纳入 Todo/reply commitment 和 required lease release,Windows 的 unattributable 又没有恢复转移,所以 public terminal 名称仍宽于实际保证。最小修复后请重跑 File/SQLite effect race、lease failure→retry、Windows CLI/MCP stop、现有 native Host suites 与完整 required CI。
我的整体评价
REQUEST_CHANGES。需求本身必要,stop sidecar、typed owner、只读 holder/Host facts、CLI/MCP 共用入口的方向也合理;当前 18 文件范围相对这个高风险能力并非单纯“代码太多”。但 exact head 没有修改上轮三个 blocker 的 owner,6 个独立负例全部复现,因此 before/after 的可观察收益仍达不到“安全停止并继续”的承诺,change proportionality 与 terminal authority 都是 not_yet_proven。建议只在现有 Todo/result/lease/Host authority 上做有界修复,不再增加平行生命周期;修复后按新 exact head 复审。
English verdict: REQUEST_CHANGES — exact head 796f688674c3a404daec89cfa6d3b5801683e292. This merge-only head leaves all three blockers unchanged and independently reproducible on File/SQLite: stop can settle after Todo/reply effects commit, failed required lease release is terminal and never retried, and launched Windows Hosts have no converging or fail-fast stop path. Focused Python/Node/static checks pass, but they do not cover these terminal-contract counterexamples.
…e and the platform Three blockers, all reproduced on the reviewed head before fixing. **A stop could settle after the member's effects committed.** The worker ran one pre-check and then committed Todo completion and reply publication outside any lock the stop takes, so a stop written first still reported `settled` for work that had landed. Both effects now commit inside the dispatch lock a stop also takes, so the two sides linearize: a stop written first means neither effect runs, and a stop written after leaves their acceptance intact. `_observe` gained a locked entry point because the kernel file lock is not reentrant and the widened critical section must not nest. **A failed required lease release was terminal and never retried.** `_settle_stop` did not pass the release to the typed decision, so an authority failure produced `settled` with `lease_released=false` on both backends and every later read returned the same receipt — leaving the member's Todo blocked until the lease TTL with resume already refused. `decideDelegationStop` now treats a required release as part of what `settled` promises, and the settle read retries the release under the stop's own lock. An operation that held no required lease omits the fact rather than claiming a release. **A launched Host on Windows could never settle.** `host_process_drain` returned `unattributable` whenever `killpg` was absent, so a finished Host left the stop pending forever with no converging or actionable path. That case is now `unsupported_platform`, distinct from an attribution failure, and `stop --execute` fails fast with an error naming the platform boundary instead of returning a receipt no read can settle. The public reference documents the lease obligation, the effect ordering and that boundary in both languages. Coverage: a stop written before the effects (both backends) asserts neither effect commits and the receipt still settles; a failed release asserts `acknowledged`/`required_lease_release_unproven` and that the next read retries and settles; a Host on a platform without process groups asserts the actionable failure. The TS decision pins the lease obligation on both the acknowledged and vanished-holder paths, mutation-checked by dropping it. Signed-off-by: song <liusongstep@gmail.com>
Re-review request — exact head
|
huangruiteng
left a comment
There was a problem hiding this comment.
Exact head: 5308@01be96c65c7c1c226ab62f99f1536e9e3979072a; immutable base: fd5f31bb3ad57448c32df6c7cddde36d07446934。本次按 LoopX PR-review capability policy 12 检查完整差异和上一评审以来的修复;未查询或等待 GitHub CI。
动机
已有 delegated member 需要可验证的停止入口,直接杀 worker 不能证明 Host 后代进程、Turn lane、Todo 和 hard lease 已妥善收尾。这是 roadmap R2 的有界 CLI/MCP 生命周期增量;它不等同于整个 App/Lark/跨宿主取消或重启验收。用户需要可信的“停止完成”,随后才能处理未完成任务或启动新 operation。当前正向 stop 有效,但持久 ACK 后的 crash 会留下活跃 lease 而读回 terminal settled,持续推进和用户恢复路径仍有缺口。
改动思路
CLI 与 host-bound MCP 共用 Delegations.stop,复用固定 binding、operation single-flight、dispatch fence、原生 task-lease authority 及 TS Host supervisor。stop sidecar 保存不可由 running 状态推导的明确停止意图;worker ACK 与 lane/Host drain 分别表达不同事实,TS owner 据此决策 requested、acknowledged、settled 或 unknown。只停止原 operation,不接管 unrelated Turn,不给外部 requester 或跨 host 新权限。host.json 是 supervisor 产生的归属事实,不是用户任意写一个 pid 就能获准杀进程。
比全新团队生命周期 framework 更小的这条边界是合理的。但 lease obligation 已有 canonical operation source,不能在恢复时仅凭 sidecar 中“缺字段”推导“不需要释放”。修复应复用这一来源,让 ACK 与释放意图在持久恢复上闭合,不增加第三套 lease authority。
具体改动
完整 18 文件覆盖 CLI stop、Python collaboration service、typed observation/stop、inventory/context/effect transport、原生 Host process owner 与 bridge/transport,两份文档及 CLI、TS、真实进程回归。相对上一评审头 796f688674c3a404daec89cfa6d3b5801683e292,本头把正常完成与 reply 发布放进 stop 的 dispatch fence,增加已知 lease 释放失败的重试,并在不支持进程组证明的平台给出明确拒绝。这些改动解决了之前的具体问题,但 lease 测试只覆盖已经持久化 {required: true, released: false} 的 sidecar,遗漏了 ACK 与该字段落盘之间的 crash。
关键代码讲解
Delegations.stop(collaboration_mcp.py:932)要求 execute、验证原 binding,持久化 intent 后信号通知持锁 worker。prepared/running/turn_returned 可停止;accepted/rejected 是无副作用 noop;stopped operation 不再 resume 或复用旧 id。_acknowledge_stop(:1039)在 operation 与 dispatch 边界下把状态改为 stopped,并先保存 ACK。之后才清 bootstrap、释放 lease、再次保存 lease 结果。它正确区分 worker/requester ACK,但这两次持久化之间存在恢复窗口。_settle_stop(:1102)读 operation lock、lane holder、Host drain,并对 sidecar 已知的 required lease 重试释放。第 1129–1139 行只看stop.lease,没有从row.task_lease.required恢复 obligation;缺值时不向 TS 传lease_released: false。decideDelegationStop(delegation.ts:367)持有 typed phase 决策,缺少lease_released被当作可结算。Python 调用端未正确区分“不需要 lease”与“释放事实尚未写下”,因此 TS 收到完整 ACK/drain 事实后返回 settled。- Host process transport 与 TS spawn supervisor 记录独立 Host 进程组,停止会等到原 Host 与同组后代退出;inventory 过滤 stop/host sidecar,context 投影 stopped。SIGTERM 没有 stop intent 时仍走原可恢复行为,而非偷偷转换成 cancel。
对主干的风险
[P1] ACK 后 crash 会把仍有活跃 hard lease 的 stop 永久标为 settled
位置:collaboration_mcp.py:1129–1139,上游写窗口是第 1070–1085 行。
独立验证在隔离的 synthetic Goal 中,通过真实 native authority acquire 取得 required lease,给原 operation 保存其实际身份与 version;只在 _release_delegation_lease 的入口注入进程退出,以模拟 ACK 已持久化、释放尚未发生。没有 mock lease 的最终状态或 settlement 结果。此时 operation 已 stopped、ACK 存在,但 stop.lease 仍是 null。恢复真实 release 实现,再从 public stop(..., execute=True) 读回,得到 settled + null lease;独立 native inspect_task_lease 却仍报告 active。真实 File 与 SQLite provider 都复现,2/2 未满足独立的终态后置条件。
_settle_stop 把 null 变成空对象,跳过 required-lease retry,再把 settled 固定成 terminal;后续读回不再修复。用户已收到“停止完成”,但原 Todo 仍被旧 lease 阻塞至 TTL。最低修复:ACK 落盘时保存不可丢失的 required lease obligation/identity,恢复时同时依据 operation 的 canonical required lease 判断,释放或独立确认释放后才允许 settled。不要把缺失 receipt 等同于无 obligation。回归须在 ACK、release 与 receipt 各持久化边界注入 process loss,重新构建 service,再验证 native lease readback;覆盖 File/SQLite,不用“mock release 返回 true”作为后置条件。
本次 62 项 Python(含 CLI/MCP、真实 worker/Host/child/lane)及 29 项相关 TS 测试通过;typecheck 与改动路径 ruff 通过。独立 source premerge 的 5 direct、5 catalog、8 risk、1 public/private boundary 均通过;但缺失 crash case 不会被这些绿灯证明正确。相同普通未请求 stop 的任务在 immutable base/head、File/SQLite 上各 2/2 通过:一次 Host、相同 accepted/canonical completion/readback、无 stop sidecar,默认路径没有被这个新增入口污染。所有测试使用该源码环境;没有 paid model call,也没有修改活跃项目状态。
语义与 CI 对齐
settled 的当前 obligation 包含实际 required lease 已释放,不只是停止了本地进程;本 PR 扩展既有 delegation observation vocabulary 和 local stop contract,不能把 machine-enforced drain/lease 条件称作 advisory。状态采用 typed phase/decision owner,不是文字或 substring 分类,通用错误仍 domain-neutral。违规是 Python producer 的未知事实被编码成“没有 required lease”,不是 enum 名称本身。复审命令须覆盖 tests/test_local_delegation.py、tests/test_delegation_cli.py、tests/control_plane/test_host_process.py、相关 TS tests,加上上述真实 authority crash-recovery 对照。此旧头的 inventory 不支持新的 changed-diff advisory 参数;支持的完整 inventory 报告及 full-tree semantic drift 均通过。CI 不参与本次判定。
我的整体评价
English verdict: REQUEST_CHANGES — exact head 01be96c65c7c1c226ab62f99f1536e9e3979072a; process loss after durable ACK produces terminal settled while a native required lease remains active, reproduced on real File/SQLite authority.
本地 stop 的目标、边界与复用方式合理,不需要重做整套架构或因为代码量拒绝。previous-review-to-head 已实质改善 effect linearization、已记录的 lease failure retry 和平台诊断;但完整 PR 的 long_horizon 与 user_experience 仍因终态失真判为 regression。面向下一次改动的有界 refine 应在现有 ACK/settle owner 收敛 lease obligation 的唯一来源,并用持久化边界测试覆盖,而不是添加互相同步的布尔旗标。真实 stop intent 与历史 operation 兼容义务应保留;未验证 App/Lark、跨宿主、PostgreSQL 或 Windows 原生执行,不据模拟平台测试声称这些验收已完成。修复上述 P1 后再复审;当前不建议合并,也没有自合并权限。
`_acknowledge_stop` persists the ACK before it releases the hard lease, so a process loss between those two writes leaves the sidecar with no `lease` field at all. `_settle_stop` read the obligation from that field, took the absence as "no required lease", and settled: the caller was told the member is safely stopped while the native authority still reported its lease active, with resume already refused and the Todo blocked until the TTL. Reproduced on both File and SQLite by acquiring a real required lease and losing the process at the release entry. The obligation now comes from the operation record — the canonical `row.task_lease.required`, which the crash cannot lose. Absent sidecar state means "the release result has not been written yet", not "nothing was owed", so the release is attempted and only an actual release lets the decision settle. That is one source of truth rather than a second boolean kept in sync with it; the released fact is still persisted beside the stop so a later read does not release twice. Coverage pins both ends of the window on both authorities: - a crash after the ACK with the lease still owed, and a release that succeeds on the next read, settles and reports `lease_released: true`; - the same window with the release still failing stays `acknowledged` with `required_lease_release_unproven` instead of reporting `settled`. Mutation-checked: deriving the obligation from the sidecar again fails all four. Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Re-review request — exact head
|
huangruiteng
left a comment
There was a problem hiding this comment.
复审结论:REQUEST_CHANGES,当前 head 8468544d0a8bd8526fa336d655701a64e74918b8。上一轮 ACK/lease 崩溃窗口的修复已验证;整份 PR 仍有一个 P1:恢复已通过验证、尚未结算的 Turn 时,Todo completion 没有进入 stop 的同一把 dispatch 锁,因此 stop 可以先落盘,Todo 随后被完成,最终仍返回 settled/stopped。
动机
现有 coordinator pause 只暂停协调者,已经启动的 delegated member 仍可能继续执行。这个 PR 让操作者经 CLI/MCP 明确停止一个 member,并区分“已请求”“已确认”“已完成清理”,保留证据,由协调者选择新的 operation 继续任务。这是有用的有界增量:单纯发 SIGTERM 无法证明原生 Host、worker、lane 和必需 lease 都已释放;不需要扩展成整个团队、跨主机或 App 的停止协议。
从普通任务看,最短流程应是:定位 operation → 显式 stop → 读取可信回执 → 检查未完成 Todo → 用新 operation 继续。恢复路径把已停止的 Todo 标成完成,会破坏最后两步;这不是仅影响诊断文本的问题。
改动思路
继续复用既有 delegation 生命周期、原生 Turn/lease authority 和 TS Host supervisor。停止意图保存于独立 .stop.json,Host attribution 保存于 .host.json;这些分别承载无法从执行状态推导的用户意图和实际进程身份。decideDelegationStop 在现有 TS owner 内根据 ACK、锁、lane、Host drain 和 lease release 事实推导 phase,Python 负责文件/进程 IO,没有另建一套 Goal completion authority。
stopped 是该 operation 的终态,旧 id 禁止 resume;停止不等于任务完成。已 accepted/rejected 的记录保持原结论,普通 SIGTERM 也不自动成为有授权的 stop。这个设计比强行终止进程或把超时当作成功更完整,新增 CLI、持久意图和 readback 的成本与问题相称。
具体改动
本轮按当前 merge base 67930ab6af78491f10ca3de4ff74ef7a39954a51 审查全部 18 个文件(+1666/-76),而非只看上一轮修复:
- CLI/MCP:新增
delegation stop --execute/stop_delegation,接入现有绑定和 operation 身份;read/wait返回 stop 状态。没有新增停止整个 Goal 或授予 peer 独立权限的入口。 - Worker/transport:stop 与 dispatch 共用 fence,启动记录传入原生 Host transport;复用 TS supervisor 的进程组清理,Python 只观察是否已 drained。外地主机、无法证明 drain、超时和 lease release 未证明都不能冒充 settled。
- TS 状态与投影:delegation transition/stop decision、effect handler、inventory、context 和 subagent context 一起识别
stopped;sidecar 不混入 operation inventory,旧 operation 的 resume 被拒绝。 - 文档与测试:两份文档解释 coordinator pause、单 member stop、清理回执和新 operation 恢复;6 个测试文件覆盖 CLI、inventory、TS transition 和 Host/worker。App/Lark UI 没有改变,本次可用入口明确是 CLI/MCP,不能据此宣称已交付 App 停止按钮。
关键路径:stop(collaboration_mcp.py:990)写停止意图;_acknowledge_stop(1097)记录 ACK 并尝试释放 lease;_settle_stop(1160)重读权威事实;TS decideDelegationStop(delegation.ts:367)决定回执 phase。普通完成路径(1663–1688)把 Todo completion、结果发布和 accepted transition 放入同一把 dispatch 锁。
相对上一轮 01be96c…,最新修复从 operation 的 task_lease.required 推导释放义务,避免 ACK 已持久化、lease 结果尚未写入时把缺字段误认为无需释放。以隔离的真实 File/SQLite authority 和 hard lease 分别注入“释放前失去进程”“释放后但结果落盘前失去进程”,重新实例化服务、公开 stop 及独立 native lease inspect,4 个用例全部通过。这个旧 blocker 已关闭;仓库新增测试本身使用 mocked release,所以本轮补了实际 authority 验证。期间合入的 workspace preflight/main 变更以当前 base 为准,不算本 PR 的新增交付。
对主干的风险
[P1] 恢复 Turn 的 completion 仍可越过已生效的 stop。 位置:恢复路径的检查与 completion。
当 turn_result 尚未 committed、但 _validated_turn_journal 已确认真实 Host 输出和 task validation 时,1626 行只检查 stop,1627 行在 dispatch 锁外调用 _complete_delegated_todo。若 stop 在这个检查之后、completion 之前先落盘,native completion 仍会执行;1629 行才看到 stop 并进入 ACK/settlement。后面的 1663 行锁无法保护已经执行过的这个 effect。
独立反例在 File 和 SQLite 两种后端均复现:使用实际 Host/task validation 输出构造持久化的可恢复 journal 边界,在恢复分支调用 completion 之前经公开 stop 写入意图,然后让原 completion 经 native CLI/authority 执行。最终公开回执是 phase=settled, status=stopped,独立读取 canonical Todo 却是 done=true。测试的独立 oracle 是“stop 先落盘则 Todo 仍打开”,两个用例均失败。这里注入了恢复 journal 和确定的 interleaving,并抑制了同步测试进程的 worker signal;没有声称用真实 OS crash 或真实信号调度复现这一竞态,也没有 mock completion 或 canonical done 后置条件。
最小修复:让恢复路径的 stop 检查与 Todo completion 也通过现有 dispatch fence 线性化,并检查后续结算/结果发布对同一 operation 的一致性。补两个后端的回归:stop 先赢时不完成 Todo、不发布成功结果,返回停止回执;completion 先赢时保留真实已提交事实,不将其改报为未完成。不要只在 completion 前后再加一次无锁检查。
验证:73 项相关 Python 测试、29 项 TS 测试、control-plane typecheck、changed Python ruff、完整 premerge(含语义词汇、维护性、输出预算及 public/private boundary)通过。相同 synthetic harness 在 immutable base/head 的 File/SQLite 普通无 stop 路径均得到 accepted、一次 Host 调用、一次结果返回、Todo done,且没有 stop projection/sidecar;当前 head 的两个 operation scope/恢复用例也通过:同 Goal 的未覆盖 sibling 保持 prepared,旧 id resume 被拒,新 id 能实际执行到 accepted。上述通过项没有覆盖 P1 的恢复 interleaving,不能抵消它。按当前 review policy 未查询或等待远端 CI。
我的整体评价
停止能力的需求、归属和整批范围成立,上一轮 lease 修复也成立。当前阻塞是同一 completion effect 有两个路径,其中恢复路径遗漏 fence,导致终态与 canonical Todo 不一致。建议在现有 owner 内统一这个 effect 的线性化入口;这也是本轮 future-facing 检查认为最有价值的有界简化,避免下一次恢复规则变更再次漏掉一条路径,不需要新框架或语言迁移。
保留的验证边界:隔离源 checkout 和 synthetic native Host,未运行付费模型、跨主机停止、PostgreSQL 或 Windows 进程组;本 PR 明示的单主机边界不因这些未测维度被扩大。本次请求修复 P1 并重新验证上述先后顺序,再对新 exact head 复审。PR 保持打开,合并留给维护者。
English verdict: REQUEST_CHANGES - head 8468544. The ACK/lease crash-window fix passes real File/SQLite lease recovery checks. However, validated-Turn recovery completes the canonical Todo outside the dispatch fence: a persisted stop can win first, yet the native completion commits and the final receipt says settled/stopped. Reproduced with an injected recoverable journal and deterministic interleaving on both backends, without mocking completion. 73 Python tests, 29 TS tests, typecheck, ruff, baseline/head no-stop parity, scoped continuation and premerge pass. Fence this recovery effect and add both ordering cases before approval.
Signed-off-by: song <liusongstep@gmail.com>
|
This pull request has merge conflicts with Choose the remote for the base repository, not an out-of-date fork. git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEADFor a same-repository clone whose Keep the DCO |
…ilure Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com> # Conflicts: # tests/test_delegation_authority_recheck.py # tests/test_delegation_preflight.py # tests/test_delegation_preview_reuse.py # tests/test_local_delegation.py
|
Published shared-fixture integration at Merged #5535's fixture repair and its regression coverage; the only tree changes from Current validation: 39 relevant stop/fixture/caller tests passed; standard premerge passed 19 selected checks plus 5 direct checks. Built the chat bundle and passed the packaged The body now names the current R1/R2/R3 proof obligations and separates passed, historical and untested evidence. The known main backpressure failure remains disclosed. New-head CI and independent review remain required; no automatic approval, review dismissal, installation or merge. |
songoow
left a comment
There was a problem hiding this comment.
Reviewer: model_agent | GPT-6 | OpenAI | self_reported
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
动机
委派任务的请求方需要停止自己获准控制的一次执行,并确认旧执行已经退出、原租约不再占用任务。此前只能撤销后续授权或手工终止进程;这个 PR 增加显式停止和可重读回执,让调用者据此决定是否可以接棒。
正常停止、租约恢复和授权范围已有实证,但内层 Host 记录丢失时仍会把活着的执行报为 settled,并释放它的租约。
本轮评审本机单操作停止、共享 Host 边界及相关展示,不要求整团队停止、跨宿主控制、Windows 原生 drain、PostgreSQL 或 Lark 的完整交付。这是作者账户的复核,我参与过上轮修复;本次重新执行验证,不把自己的修复说明或旧测试数量当成独立批准。
评审 head:c49c1fdbcc5a243dffd2465c583b2696529453fe;完整 diff merge-base:99839aeb8fed5fae38a5d319391cd050672a6508。结论 REQUEST_CHANGES:一项新确认的 P1 恢复边界缺陷;前一 head 的 CI 失败归因与当前 head 的待完成检查分开列出。
改动思路
可用的停止需要同时保持原操作身份、拒绝迟到效果并判断执行是否退出。只发信号不能给出这些保证;把 #5534 的“先撤销、后观察退出”直接换进来又会改变当前 settled 契约。因此仍应遵循最新 R1–R3 评审明确的三个既有 owner:canonical authority 管租约与提交资格,Host 管完整执行资源,delegation 管停止意图、编排和真实回执。
当前实现的有界收敛值得保留:TS 的 stop step 已区分动作许可和最终回执,canonical 查询不再被空/过期注解绕过,嵌套记录布局和平台限制已经收回 Host transport。问题不是需要另建服务,而是该 Host 观察仍把不完整来源当成已证明的缺席。继续增加 delegation 侧的 lease/拓扑猜测会重新破坏这次职责整理。
具体改动
完整 diff 33 文件,+3032/−171:运行时/产品 17 文件 +1122/−132,测试 13 文件 +1757/−32,文档和浏览器场景 3 文件 +153/−7。已读完整变更及相邻调用方。发布前检测到 head 从 273363daa 更新:新增仅为四个测试文件的 fixture 清理、失败 setup 回收验证及格式调整。已重新生成评审包、复读差异并重跑受影响检查;运行时、文档、前端、TS 测试和原有 stop/lease 测试源码与前一 head 一致。旧运行保留原版本,不改名为新 head 结果。
关键代码讲解
Delegations.stop/_settle_stop:CLI/MCP 共用精确 binding 校验、独立 stop intent 和稳定 stop_id。_execute在 dispatch fence 内覆盖 Todo completion、原 Turn settlement 和结果发布;resume 拒绝已经停止的操作。decideDelegationStop:返回resolve_lease或record的判别联合,要求显式 lease fact;unchecked 不能被省略成“不需要”,时间本身不产生 settled。delegation_stop_lease.release:按原 owner/key/可用 epoch 查询 canonical lease,再以当前版本释放;released/not_owed 与两类未证明状态分开。真实 File/SQLite 的空、过期、缺失、畸形注解与替换 generation 测试通过。execution_host_drain:集中普通/嵌套执行观察与平台政策;private CLI 传递记录地址,transport 在用户 Host 前消费 marker。TS spawned/openInput hooks 和最新已证明 expiry 的生命周期保持原 owner。这里的缺记录分支有下述 P1。delegationState与 GoalTeamWork:stopped 以“停止已登记”及独立统计桶展示;context/inventory 传递状态并排除 sidecar。没有新增停止按钮、Lark 控件或配置面。打包浏览器测试确认标签没有宣称“执行已释放”。
其他变更包括 CLI --execute 门槛、MCP stop 工具、signal unwinding、重试/恢复文档和私有 fixture runtime 回收;不是新的 Agent 管理权限或第二份租约存储。
修改前契约:docs/reference/local-delegation.md,固定版本 58574643c0717508ec7cb3e2c35415566bbf720c:
| 原契约章节 | 本 head 判断 |
|---|---|
| Activate | implemented(有界实证):同一 requester 的精确 grant;旁路记录不变,未来新 operation 不被自动停止,撤销 grant 在 stop 写入前拒绝,恢复后实际执行到 accepted。 |
| Acceptance and return | implemented(本次相关路径):canonical completion 仍在既有 owner;stop/恢复 completion 的两种排序及原 Turn 结算测试通过,不宣称 Goal 完成。 |
| Disconnect and recovery | not_met:本次记录丢失反例可得到不真实的最终停止回执,不能安全据此进行后续接棒。 |
对最新评审:R1 已验证修复;R3 已验证修复;R2 的职责移动已完成,但其“完整执行观察、保留缺失归属保护”仍未满足。 不恢复维护者已经移除的“每次历史失败都必须先 RCA”要求。
对主干的风险
[P1] 内层 Host 记录缺失会让仍活着的嵌套执行被结算,并提前释放原租约。
host_process_drain:94–95 把 FileNotFoundError 返回为 not_launched;execution_host_drain:134–137 把它与外层 drained 合并成 drained。随后 _settle_stop 获得 resolve_lease、执行 canonical release,并持久化 settled。终态回执会被直接重放,不会在下次读取时重新证明 Host 已退出。
我在当前 c49c1fdb 的隔离 File/SQLite 上重新执行既有真实 nested-stop 场景的记录丢失扩展:启动真实 hard-lease worker、原生 Host 和忽略 TERM 的同组子进程;只暂停该 fixture 的内层 supervisor。先调用真实 stop,确认它保持 acknowledged、canonical lease 仍 active、实际 Host/子进程仍活跃且外层组已经退出。然后仅模拟丢失本次执行的内层归属记录,再以同一 operation 调 stop:
| 独立读回 | File | SQLite |
|---|---|---|
| 记录完整时的 phase | acknowledged | acknowledged |
| 内层记录缺失后的 phase | settled | settled |
| stop_id | 与第一次相同 | 与第一次相同 |
| canonical lease | released | released |
| 实际 Host / 子进程 | 都仍活着 | 都仍活着 |
两例均失败在“活执行不能 settled”的断言。没有 mock stop、ACK、租约读写或 OS 退出事实;只故障注入记录丢失和 supervisor 暂停,finally 恢复记录并清理自己创建的进程。不声称正常流程已经产生过记录丢失,也不声称线上发生过事故。 这是当前恢复承诺的反例。
现有 test_execution_drain_needs_every_group_a_leased_run_launched:383 恰把 outer 已退出、inner 缺失编码成期望 drained,因此已有测试通过不能反驳它;另一个“missing or unreadable”测试覆盖的是旧格式/损坏 JSON,没有覆盖实际 inner 文件丢失。
最小修复:让 Host owner 区分“有正面证据证明从未启动”和“预期的内层记录缺失”。已知 leased/nested 执行不能仅凭缺文件报告 not_launched/drained;保留 acknowledged/unattributable 和租约,待原执行的完整证据恢复后再结算。修正现有测试中 outer-drained + missing-inner = drained 的期望,并把真实记录丢失反例并入既有 File/SQLite nested-stop 用例。无需新增公共 phase、全量 TS 迁移、第二份 journal 服务或把 #5534 的替代契约混入这里。
本轮实际验证(保留源版本):
-
当前
c49c1fdb的独立组合运行:记录丢失 2 failed、范围及 fixture 隔离 4 passed;新增 setup-failure 清理测试另 1 passed。前一 head 的记录丢失两例也失败,未用这些旧失败替代新 head 的复现。 -
前一 head
273363daa上 Stop/recovery/nested lease/真实 CLI 的 38 项通过;另有原生 Host 12 项通过(main 相同公共用例 7 项通过)、平台副作用前拒绝 4 项通过、inventory 6 项通过。这些运行覆盖有重叠,不相加成“独立总数”。 -
新范围探针在 File/SQLite 2 项通过(已在当前 head 重跑,包含于上述 4 passed):目标停止、邻近操作 bytes 不变、未来新操作到 accepted 且 Host 恰一次、撤销 grant 无写入、恢复配置后继续、已 accepted 的 stop 为 noop。初版探针把既有 EffectRuntimeRejected 错按 ValueError 捕获,纠正探针后跑通;没有修改产品代码或放宽授权断言。
-
273363daa上 44 项 TS 测试、control-plane typecheck、CI 配置的 mypy(19 文件)、所选变更 Python Ruff 通过。 当前 head 对应生产源码及这些 TS 用例未变,作为有明确失效核对的复用证据;不是重新运行的声明。 普通 Host 的五个 base/head 配对场景(输入输出、继承环境、显式环境排除 ambient、超时、非零退出)返回对象/stdout/stderr 完全一致,无语义归一化;场景脚本 SHA256377d8d6c91881a41a12f960fdcd90321471cc7f7e2e267f2bb67daefe4a7f993。 -
273363daa上 Dashboard/Chat 构建及 packagedteam-evidence场景通过;当前 head 的 renderer、场景及构建入口未变,复用该证据。检查了桌面视口,并覆盖停止标签、移动布局和读回不可用/重试。浏览器使用合成 API 状态,证明 renderer 的表达和交互;实际 stop/lease 来自上述独立 CLI/服务测试,不冒充已安装 App 或浏览器直连真实停止后端验收。 -
新增记录丢失反例 2 failed,如上表;这才是本轮新增阻塞的直接依据。公共边界扫描 33 路径通过,当前 58 个 PR 提交都有 DCO trailer。
CI 与基线归因: 已检查前一 head 273363daa 的运行 37163168073,该运行四个 pytest shard 失败,pytest/merge-gate 汇总也失败;较早运行的取消项不当成新的失败。三个抽样失败在未经修改的 main 99839aeb8fed5fae38a5d319391cd050672a6508 与前一 head 273363daa 的同命令、同 node ID、完整 failure message 均一致:goal-instance-binding inventory 缺 codex_cli_session_binding;late prior-instance result 的 recipient 授权错误;monitor quiet-due 的 selected_todo KeyError。这三例归为 pre_existing_unrelated;没有据此推断其余几十例都与 PR 无关,剩余归因保留 unresolved。当前 c49c1fdb 的新 CI 仍有 pending 检查,尚未获资格;没有把前一 head 的红色检查冒充当前新运行结果。
273363daa 的全树 docs-governance 同样失败于 main 已有的 external-evidence-research RFC 中英两处 dated-heading 结构违规;本 PR 不改那些文件,已有相同基线命令与失败签名。没有把它算成本轮新增 stop 缺陷,也没有把文档检查说成通过。未重跑全仓 Python、付费模型、Windows 原生 drain、PostgreSQL 或跨宿主路径。
语义与 CI 对齐
开发期 advisory 对完整 PR diff 检出三个本地 stop 分类集合;它不分析 TS union,不能作为完整语义证明。新 stop phases 与 lease facts 属于既有 delegation typed owner,Host 归属字段属于其私有记录;不应借相同单词误复用其他 scheduler/replay 的 settled 语义。没有 substring denylist 或领域专用控制面义务,--execute、binding、dispatch fence 都是强制检查而非 guidance。
真正的语义错误是“缺观察”被分类成“从未启动”,而不是最终 TS phase 计算缺一个布尔值。新 stop 工具在已配置 delegation 服务中可见是显式新增能力;没有宣称整个工具集在未调用 stop 时逐字不变。无 stop 调用的共享原生 Host 输入/结果配对保持一致,未来 operation 也不继承旧停止意图。
我的整体评价
REQUEST_CHANGES。 正常路径的进展和 R1/R3 收敛均有当前证据,33 文件的主题仍集中在单操作停止、证明与回读;没有理由因测试行数多就删掉独立失败模式。但当前 long_horizon 与 user_experience 仍有明确 regression:一次不完整恢复读回就能释放仍在执行的持有者租约,并给调用者不可再证明的成功回执。正确的“停止已登记”标签无法弥补后端 settled 的错误。
面向后续修改的检查已落到一个有界问题:修复 Host owner 的来源完整性和真实 never-launched 证据,修改错误测试 oracle,保留 TS 行为决定、canonical lease 和 worker 编排的现有分工。不能只靠回执字段或 helper 搬家宣称 owner 已收敛;也不需要新建三个服务。#5534 仍是待选择的替代方案,不是本 PR 的验收替代物。
先修复该 P1,并保留当前 CI 的明确归因范围后再复审。没有撤销维护者意见、改动运行时代码、合并 PR 或声称父级协作目标完成。
English verdict: REQUEST_CHANGES - c49c1fd. R1 canonical lease recovery and R3 typed action/receipt separation are verified, but missing inner Host attribution can turn an acknowledged live execution into settled and release its real lease on both File and SQLite. Preserve unknown/missing drain evidence in the Host owner and repair the test oracle. Focused stop, native Host, scope, TS and packaged-renderer checks pass; the two real record-loss counterexamples fail. Three earlier-head CI samples match main; other earlier failures remain unattributed, and current-head CI is pending. No merge or whole-platform qualification.
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
Published the bounded repair for the current P1 finding at
Validation on the repaired tree: 42 passed, including File/SQLite record-loss/recovery and real native Host paths; 8 renewal-fault cases excluded from this repair run (they passed earlier in this session). Standard premerge against immutable main Four existing files, +48/-10, split into signed runtime and regression commits. Existing frontend/Lark behavior is unchanged: the repair affects private Host evidence consumed by the existing stop receipt. Prior-head Windows runtime-shutdown CI remains unattributed; the new head needs its own CI and review. This repair response is not an independent approval, review dismissal or merge. |
Signed-off-by: song <liusongstep@gmail.com>
|
Follow-up repair at The current CI investigation found a concrete shared preview-supervisor race: cancellation clears timers, but a suspended result callback can resume and arm a new idle timer, leaving the bridge alive after its worker exits. A deterministic real-process regression forces that ordering. It fails before the repair, then passes with the fix. Validation: full preview suite 39 passed; isolated Linux retirement/backpressure/cancellation selection 5 passed; Host/preview TS 22 passed; standard risk-based premerge and Ruff passed. The earlier missing-inner-Host-record P1 repair remains in this head: missing expected attribution cannot establish drain or release the original lease. Review resolution still requires current-head assessment; this comment does not dismiss any review. CI on the preceding head had 72 failures. 69 have identical testcase/failure detail on immutable main |
|
Current-implementation finding at
The reproducer uses the actual leased Host transport, canonical lease and public The repair in progress keeps ownership in the Host transport: either surviving record must retain whether a leased peer is required, and missing peer evidence must remain unproven. Validation must include record loss on both sides, real never-launched cleanup, active-lease preservation, restoration followed by actual exit and same-stop-id settlement. No new lifecycle service or timeout increase is needed. The previous missing-inner-record repair and green focused suites did not establish this reverse direction. Current implementation approval remains withheld until this repair is qualified. |
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…ution Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
… lost Signed-off-by: song <liusongstep@gmail.com>
|
Implementation repair published at The deeper review found that fixing only a missing inner record was insufficient: losing the outer record, or both records, could still settle a live leased execution. Both cases were reproduced through real File/SQLite authority and the actual leased transport/public stop path. The Host owner now requires positive non-execution evidence created before a new operation can launch. Either surviving leased record requires the other; missing primary proof is unproven. Replaying an existing operation never recreates missing proof. Tests retain the active original lease while execution lives, restore the actual evidence, let the owned process exit, then verify settlement with the same stop id. The test-only single-record drain facade was retired in favor of the production aggregate entrypoint. No additional stop phase, lease authority, lifecycle service or timeout increase was added. Final source validation: 115 passed across delegation/Host/recovery/CLI/inventory; 2 passed File/SQLite scope and useful recovery; 19 selected + 5 direct premerge checks passed; Ruff and diff hygiene passed. New private supervision markers are consumed before user Host launch. The interruption of one intermediate test run is not counted as qualification; the final 115-case run was completed after the implementation correction. This closes the reproduced implementation counterexamples. The complete current-head review record and CI observation are being reconciled separately; this repair comment is not a maintainer approval or merge authorization. |
|
Current-head qualification update for The reproduced delegation implementation defects are repaired. All 379 delegation/Host cases passed in the full CI run; the real File/SQLite and isolated Linux stop checks also passed. Seventy failing cases were independently matched to immutable main by identity and failure detail, including controlled same-input checks for the runtime-fingerprint differences. One additional current failure remains unresolved: Standalone, full-module and targeted concurrent controls pass, but those do not erase the repeated full-shard failure. The remaining evidence gap is to capture its underlying exception under the original collection/sharding conditions, then repair or independently attribute the responsible path. No acceptance check has been weakened, no final APPROVE has been published for this head, and no merge is claimed. |
Signed-off-by: song <liusongstep@gmail.com>
…-20261004 Signed-off-by: song <liusongstep@gmail.com>
|
Current implementation finding at In isolated real File and SQLite fixtures, an operation acquired its original canonical hard lease before launch. After removing its optional
This is an additional current implementation blocker, separate from the notification CI diagnosis. It must be repaired and verified before approval; prior local successes do not close it. A second R1 counterexample also reproduced on both real providers: a |
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…-20261004 Signed-off-by: song <liusongstep@gmail.com>
修复读回:7aa65e0aaaaabe889c3c3b6c6f1fe163ff4ef93d本次复查覆盖实现语义及真实进程/存储边界;不是仅修 CI。目前已确认的实现反例均已修复并推送,当前 head 的最终评审仍等待必要验证完成。这条记录不是 APPROVE,也不撤销维护者的 review。
当前验证: 整合 main 对先前失败做了新的 main/head 同选择双 worker 对照:71 项中,两侧均为相同 70 个 node ID 失败、notification 目标通过。68 个 failure-message 属性完全一致;另两个是 checkout 路径及输出截断差异,缺少 settlement binding 的断言相同。它只证明该选择的当前基线对照,不能代替完整 CI。此前 notification 在完整 CI 中失败两次,随后原完整并发 shard 通过;原因仍未知,不声称已修复该偶发失败。 剩余关闭条件: 当前 head 的 完整 Python CI 仍 pending,其他必需检查也未全部结束。必须读回当前运行,并对新增/无法归因的失败给出受影响不变量的证明后,才能完成最终 APPROVE 判断;不会要求重建全部已丢失历史现场。当前是 author-owned PR,即使最终结论为 APPROVE,也只能由本账号发布 COMMENTED 自评,GitHub 正式批准仍需独立维护者。未合并、未撤销旧 review。 |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent · gpt-6.1-sol · OpenAI · runtime_reported · xhigh
APPROVE:当前有界的单操作停止切片没有独立复现的阻塞代码问题。 精确 head 7aa65e0aaaaabe889c3c3b6c6f1fe163ff4ef93d;完整34文件差异 +3301/-173。默认 premerge 的一项既有失败和其他 review/合并权限仍需分别处理,不能把这份批准当作已具备合并资格。
动机
通过 CLI 或 MCP 委派本机成员任务的操作者需要安全终止已经委派的任务。成员任务仍在执行时,操作者需要终止指定操作;过去只能中断进程,留下执行或租约状态不确定,下一次任务可能重叠。现在可按原 operation_id 请求停止,等全部已启动进程退出后再释放其原租约;未证明退出时显示中间状态并保留重试路径。当前精确 head 在真实 CLI、MCP、File/SQLite 权威存储和实际子进程中证明了指定操作停止、租约有序释放及后续新操作正常完成。本次不认证跨主机/Windows 停止、真实模型长时间运行、整队或 Goal 完结,也不宣称安装中的应用已升级。前端当前只显示停止状态,尚无停止操作交互;Lark 操作入口、安装版应用和更广平台资格仍未在这份有界切片中交付。
改动思路
停止意图、原生进程退出和任务租约是三种不同事实。这里复用原 operation 的 dispatch 锁阻止迟到的 Todo、Turn 结算和回复,复用 Host supervisor 管理全部已启动进程,再由现有 TypeScript 规则决定是否可以释放原租约、写最终收据。杀一个进程或先释放租约都不能单独解决这项问题。
独立检查了当前维护者 R1/R2/R3 验收框架,spec_revision:bdbd44076ede5f711a875832e5963d5dee19a6d3,以及其引用的 既有 local-delegation 契约。R1、R2、R3 均 implemented;下文给出对应路径和反例。后续 #5534 的 lease-first 方案仍是待维护者决定的提案,不能替代当前 drain-first 验收依据。
具体改动
关键函数讲解
Delegations.stop(collaboration_mcp.py:1017)从原操作解析当前 grant 和不变 binding,在 dispatch 锁内登记显式停止。CLI 缺少--execute会拒绝;MCPstop_delegation也进入同一规则。它不把普通 SIGTERM 当停止意图,也不会改写已经 accepted/rejected 的结果。decideDelegationStop(delegation.ts:394)负责 R3:锁、Turn lane、ACK、全部 Host 退出和租约事实齐备后才允许最终结算;先返回resolve_lease中间步骤,再判断最终收据。unproven不能被解释成没有义务。execution_host_drain(host_process_transport.py:140)与 TSrunLeasedHostProcess负责 R2:观察原生及嵌套 Host 进程组,缺失记录但可能已经启动时保留未知状态。真实忽略 TERM 的子进程、暂停 supervisor、丢失 inner/outer/两类记录和同 stop_id 恢复均覆盖;不会凭超时或失联宣称退出。delegation_stop_lease.release(delegation_stop_lease.py:65)负责 R1:从原操作身份向 canonical owner 核实义务,用匹配 owner/key、有效执行 epoch 和最新 version CAS 释放。真正的租约状态优先于冗余 annotation,后来取得的其他代际租约不会被释放。
正例经真实 CLI/MCP 请求停止、完整 drain、canonical readback 再启动新操作完成。独立 MCP 反例在 File、SQLite 都证明:无 grant 的 requester 被拒绝;同容器 sibling 字节不变;配置修改不能重定向原操作;停止后新增操作 accepted,未继承停止记录,两次任务合计恰好两次 Host 调用。
最强回归反例使用相同真实 native lease 和 task_lease={}:旧 bdbd44076ede5f711a875832e5963d5dee19a6d3 两种后端都显示 settled,却保留原租约 active;当前 head 两种后端都在同 execution epoch 原租约 released 后才显示 settled。这里不是从新 helper 的输出推导 oracle,也没有使用活跃业务 Goal 测试。
对主干的风险
本 head 独立完成321项 Python(179主路径 +142相邻路径)、48项 TS、typecheck、20个改动 Python 文件 Ruff、whitespace、semantic advisory 后完整 vocabulary smoke,以及实际 CLI/MCP/子进程/权威存储验证。普通 Host 输入/env、非零退出+stderr、空输出和 timeout 四组 immutable 6a172f699653a11c5a4dad578e244fe535013c4c/head 对照的完整结果、stdout/stderr 相同,无字段裁剪或错误文字规范化。200个并发成员 Turn 和双向 dispatch race 覆盖恢复与无重复效果。
打包页面的 desktop/390px team-evidence 场景通过,停止显示“已登记停止”,不冒称执行已释放;修改、重载与503不可用后的状态读回也通过。已查看整个 viewport。页面 API 是合成 fixture,不能证明真实后端排序、新鲜度或安装版采用;真实停止 owner 的证据来自上述独立 CLI/MCP/存储。界面尚无停止操作按钮,此伴随入口缺口保留。
native premerge:5项直接检查通过,19项选定检查执行,18通过、1失败。失败是输出预算的六个 review_packet_handoff_only 行:crowded/multi_agent/small 的 JSON/Markdown;同命令在 immutable 6a172f699653a11c5a4dad578e244fe535013c4c 与 head 的失败身份和全部增长/限额细节完全一致,相关 packet/budget 因果文件未改。默认检查使用较新的 origin/main;固定实际 PR 基础版本的同 workload 差分检查通过。故仅将这一具体失败归为 pre_existing_unrelated,仍保留默认 premerge 失败与单独合并 hold;没有提高限额、移除断言或宣称 premerge 全绿。
历史 notification CI 失败的原因仍未知;本次完整相邻模块与其他主路径并行验证通过并检查了实际 runtime 隔离,但不宣称历史失败已修好或整套 CI 已归因。按当前 review 配置没有获取、轮询或等待远端 CI。Pydantic 的 lifespan forward-reference warning 保留;没有断言失败。Windows、跨主机、付费模型长跑、安装版 App 和 Lark stop 效果未测,不在这次批准结论内。
我的整体评价
APPROVE,交付判定 justified_increment。 原始问题的长期价值成立:停止后能够继续真实新工作,原 operation 的恢复不重复启动,canonical 租约不会被缺失注解抹去。long_horizon improved;user_experience accepted_tradeoff:原 operation_id 与显式 effect consent 是必要边界,CLI/MCP 可用且状态真实,但 dashboard/Lark 操作交互仍由原 collaboration capability owner 后续交付,不能称整个能力已完成。
规模重新从原问题评估,而非把消除旧 finding 当作必然批准轨迹。最有价值的未来变更整理已在当前切片应用:native lease 义务、Host 拓扑事实和 typed stop 最终性各有一个现存 owner;Python 保留 IO/OS 适配,没有重建第二份状态政策。不需要另建通用 Actor 生命周期框架或迁移所有历史接口。遗漏/损坏事实、外代际和普通信号的反例同时覆盖了放行和误阻塞两个方向。
没有新的阻塞代码 finding。最强剩余资格缺口是更广平台/安装版与前端操作交互;最具体的当前质量 hold 是上述默认输出预算失败。发表并读回此 unchanged-head 批准后,仍应运行 native approval-closeout,保留未获授权协调的其他 review,不自动合并。
English verdict: APPROVE — exact head 7aa65e0. The bounded CLI/MCP operation stop independently passes321 Python/48 TS checks, real File/SQLite lease and native nested-process negatives, exact-scope MCP recovery, full ordinary-Host base/head parity and packaged display checks. R1/R2/R3 are resolved at this head. One default premerge budget failure is independently identical at the immutable base and remains a separate merge hold. No CI queried; broader UI operation, installed/platform and live-model qualification are not claimed.
|
Post-merge review reconciliation at exact head The findings in the cocolord review 5366235560 and maintainer review 5401677786, including R1's inline comment, were individually checked against the unchanged head and independently validated. Resolution evidence is in the exact-head approval:
The PR was merged while reconciliation was being checked. Native dismissal requests, including explicit |
Problem and outcome
An authorized requester needs to stop one delegated operation and know when its execution and original lease have been released. CLI
delegation stop --executeand MCPstop_delegationshare the existing binding and dispatch fence.settledrequires acknowledgement, released operation/lane holders, complete attributed Host-group drain, and a canonical lease result. The dashboard reports “Stop recorded”, which does not establish resource release or Goal completion.Current head:
7aa65e0aaaaabe889c3c3b6c6f1fe163ff4ef93d; integrated main:6a172f699653a11c5a4dad578e244fe535013c4c.Implementation and review repairs
required:false; owner/key/epoch and current-version CAS remain enforced.resolve_leasefrom finalrecord; Python supplies facts and performs protected IO. No new lifecycle service, provider or parallel lease writer is introduced.These address the R1–R3 frame, missing inner record finding, and the current-implementation record-loss finding. Stop intent remains explicit; Host attribution is private infrastructure created for ordinary delegation execution as well.
Validation
7aa65e0a, after the shared-source-grant main integration: 119 passed across local delegation, nested stop, recovery, native Host, CLI and inventory; 2 passed real File/SQLite scope-and-recovery probes. These cover normal operation, competing effects, lost records on either/both sides, missing authority markers, retained original leases, replay without fabricated proof, and same-stop-id settlement after restored evidence and actual exit.6a172f69: 19 selected + 5 direct checks passed. The integrated shared-source-grant TypeScript suite also passed 8 tests. Ruff and diff hygiene passed. The extra strict Host-transport mypy invocation retains the same three existing errors; it is not reported as passing.Boundaries and remaining gates
Only attributed POSIX process groups are proved drained. No cross-host control, native Windows complete-drain guarantee, arbitrary external-effect rollback or whole-Goal settlement is claimed. Unsupported launched execution is rejected before stop mutation. CLI/MCP expose stop; the existing dashboard reads its status; Lark has no new stop command.
#5535's fixture isolation repair is incorporated and can merge independently. #5534 is merged as a proposed, still-unselected revoke-first design alternative, not a prerequisite or replacement acceptance. Stop does not block the separate SQLite adoption/legacy-writer retirement program.
New-head review and CI remain required. At
f02de776, 379 delegation/Host cases passed in the full parallel CI run. Seventy failures were independently attributed to unchanged main; one notification-recovery case failed in two CI attempts. Its next complete original two-worker shard atf5d8606fpassed after adding an exception diagnostic; the initial cause remains unknown, and this is not called a causal repair. The current head preserves its original exception traceback in the successful synthetic-delivery test. This diagnostic passed all 10 module cases locally after integration and an intentional exception probe confirmed the traceback is retained. It does not change production error disclosure or establish that the CI failure is fixed. Mandatory checks must not be bypassed. The integrated-head delegation suite and premerge passed against main6a172f69; current-head CI is still pending. A fresh identical 71-case, two-worker selection on main6a172f69and head7aa65e0aproduced the same 70 failed node IDs and one passing notification-recovery case on each. Sixty-eight failure-message attributes match exactly; the remaining two contain different checkout paths/truncation but the same missing-settlement-binding assertion. This selected comparison does not replace the pending full CI run or establish the original intermittent notification failure cause. Implementation review and GitHub merge readiness are separate. No review dismissal, installation or merge is claimed.