Skip to content

feat(scheduler): report a host delivery window for bound app lanes - #6129

Merged
huangruiteng merged 3 commits into
loopx-project:mainfrom
GZY-SUPER-HACKER:feat/host-delivery-window
Oct 11, 2026
Merged

huangruiteng merged 3 commits into
loopx-project:mainfrom
GZY-SUPER-HACKER:feat/host-delivery-window

Conversation

@GZY-SUPER-HACKER

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Outcome basis / optional anchor: no public issue, roadmap card or RFC id is invented for this mechanism. The contributor board row that asks for the outcome is GH-A02 (S4/S10 · [Bug]: Codex App automation integration cannot detect missing scheduled prompt delivery #3927), whose status is Needs design; the slice implemented here is the one I proposed on that issue, and this PR does not claim the row was assigned to me. A maintainer may prefer a different shape and I will re-scope.

  • Goal/source and gap: a bound Codex App lane can stop receiving scheduled turns while its goal keeps reading as runnable. What loopX holds today is configuration evidence: the installed automation manifest, its ACK, and host_match_observed all describe what loopX configured, not what the host did. Every receipt check that does exist is keyed on a turn_instance_id that the host produces (heartbeatReceiptStatus, unsettled_host_turn_recovery, _record_automatic_heartbeat_stall), so when an automation never fires there is no identity to key on, those checks report "not applicable", and nothing surfaces the silence.

  • Observable before → after, with the validation row that proves it: the status projection's bound-thread rows carried only the observed host state; they now also carry delivery_window with state (fresh / stale / missing / unknown), last_observed_at, age_seconds, age_hours, and the cadence it was compared against (validation rows 1–3).

  • Issue/task and intended base: Related to [Bug]: Codex App automation integration cannot detect missing scheduled prompt delivery #3927 — this does not close it and does not name a root cause for the symptom reported there. Intended base main.

Author Declaration

  • Written by: model_agent — Claude (Anthropic), working from the operator's session and environment.

Implemented against

  • Specification and revision: no written specification; the request in this PR is the basis. The invariants it relies on were read from the current tree: loopx/control_plane/agents/host_thread_activity.py (readback resolves a goal's threads through thread_agent_bindings; observers never write to the host and report unknown rather than an optimistic state), loopx/codex_app_thread_activity.py (the Codex local-store adapter, consumed by loopx/chat_status_api.py), and loopx/doctor.py::add_promotion_readiness_freshness (the existing fresh / stale / missing / unknown freshness shape with the age exposed).
  • Criteria:
Criterion (spec clause) Disposition Symbol / path Test or command
A bound lane's delivery window is reported with the last observed tick and its age implemented host_thread_activity.py::build_host_delivery_window tests/test_host_thread_activity.py::test_delivery_window_separates_fresh_from_stale_by_the_expected_cadence
Absence of evidence is never reported as healthy implemented same (missing / unknown, never fresh) tests/test_host_thread_activity.py::test_delivery_window_never_reports_absence_as_fresh
The window is computed once, in the projection that already observed the host implemented host_thread_activity.py::attach_host_delivery_windows tests/test_host_thread_activity.py::test_attach_reports_the_window_from_an_installed_automation
The comparison carries no host-specific knowledge implemented HostDeliveryExpectation / HostDeliveryExpectationProvider are provider-agnostic; only codex_delivery_expectations is Codex-specific same tests
The expected cadence comes from what the host will actually fire implemented codex_app_thread_activity.py::_installed_automation_intervals reads the installed manifest tests/test_host_thread_activity.py::test_delivery_expectations_resolve_only_installed_active_app_automations
Providers for hosts other than Codex out_of_scope — Only Codex has a local-store observer today (codex_thread_observers is the sole HostThreadObserver implementation), and its store shapes are not a public contract. A second host needs its own observer first; writing one from guesswork would be an unverifiable adapter.
  • Self-check before submission: I read the observer, the projection, the status route, the freshness idiom in doctor.py, and the automation enumeration in control_plane/heartbeat/automation_upgrade.py before writing. I verified the reduction end to end through the real main() status route with the Codex home redirected to a disposable directory. I deliberately did not touch scheduler/heartbeat_followup.ts or the receipt freshness fence, and I did not add the window to loopx doctor's required checks — doctor's exit code is the conjunction of its required checks, so a quiet goal there would report a healthy installation as broken. The head is based on current main.

Scope And Continuation

  • Completed scope and remaining work: complete within this scope for the Codex app surface. Two decisions I took rather than leaving open, both reversible:
    • Expected-window source. The installed automation's rrule (what the host will actually fire) rather than loopX's ACK record (applied_rrule). The ACK record is only reachable through the serving effect runtime, which would give a status request a runtime dependency and a timeout it does not have today; the manifest is a plain file read, the same kind the adapter already does, and it tracks the host's real schedule if the automation is edited outside loopX. The chosen source is named in every payload (source), so a reader can tell which cadence an answer came from.
    • Tolerance. The window is interval × 2, so a lane is stale only after it misses more than one fire. A single missed fire is not yet reported.
  • Unrelated but visible: _automation_rrule_interval_minutes reads the INTERVAL= field of a minutely rrule. The canonical parse lives in TypeScript (scheduler/state_store.ts) and is only reachable through the scheduler runtime, so this is a deliberate local read of the same field, restricted to FREQ=MINUTELY; anything it cannot read reports no interval, which becomes unknown rather than a guess.
  • Slice boundary / successor: independently testable and reversible — the change only adds a field to an existing read-only projection, and one existing test asserting the exact thread payload was updated to include it. No successor is required. It does not overlap feat(codex-app): add opt-in heartbeat delivery diagnostics #4140, which is a manual opt-in canary needing a caller-asserted turn id.

Validation

  • Tested revision: e4658d8e1cf5eff07b57fc8ca7eee75df319e0e2
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
unit passed tests/test_host_thread_activity.py — 44 passed, including the window, provider and rrule cases
integration passed Every test module referencing chat_status_api, codex_app_thread_activity or host_thread_activity — 172 passed
regression_parity passed With the status-route wiring reverted, test_app_status_route_attaches_codex_thread_activity fails; with it, the module passes. The updated assertion is the one that pins the new field.
static passed ruff check clean on all four files; git diff --check clean
real_entrypoint passed The projection was produced through the real status route with the Codex home redirected to a disposable directory and no real store touched
real_backend not_applicable No live Codex App automation was installed or triggered, and the window reads only local files
  • Coverage and gaps: named plainly. (1) The window is proved against synthetic stores; a real automation that stopped firing has not been observed, and the loopX-side consequence is all that is claimed. (2) A pre-existing failure in tests/test_chat_server_cors.py::test_chat_preview_preserves_canonical_actor_refusal[sqlite] appears identically with and without this change (a Node bridge error raised from tests/control_plane/canonical_authority_fixture.py), so it is environment and not addressed here. (3) The 2× tolerance is a judgement, not a qualified value; the board row's own exit mentions a missed-delivery window without naming one.

Frontend / Visual Evidence

  • UI impact: none — the field is added to the status projection JSON and is not rendered by any current frontend.

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

  • Direction / acceptance reference, when applicable: S4 (host parity) and S10 (reliability diagnostics) are the streams the board row cites. The comparison is deliberately host-agnostic so a later provider can extend it to another surface without changing the projection, which is the S4 parity sentence.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A
  • Semantic dimensions changed, or reviewed no-impact rationale: N/A
  • Provider conformance arms run: N/A
  • Read-only legacy/file/PostgreSQL three-arm rehearsal: N/A

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths (including .loopx/, .codex/goals/, and live ACTIVE_GOAL_STATE.md).
  • I did not duplicate maintainer-owned benchmark work unless a maintainer split out a public issue for it.
  • I kept the change scoped to the linked issue/task.
  • I completed the visual evidence section for UI changes, or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer (git commit -s).

A bound Codex App lane can stop receiving scheduled turns while its goal keeps
reading as runnable. The installed automation and its ACK are configuration
evidence, and every receipt check loopX holds is keyed on a turn identity the
host produces, so none of them answer "has this lane moved lately".

Add a read-only delivery window to the existing status projection. It compares a
lane's own last observed host activity against the cadence its installed
automation will actually fire at, and reports fresh / stale / missing / unknown
with the observation age. Absence of evidence is never fresh: an unobserved lane
is missing, and a lane whose cadence cannot be read is unknown.

The comparison is host-agnostic; only the expectation provider is Codex-specific,
and it reads the automation manifest rather than the scheduler runtime so that a
status request keeps no dependency on that runtime.

Related to loopx-project#3927. Does not close it and does not name a root cause.

Signed-off-by: GZY-SUPER-HACKER <162807803+GZY-SUPER-HACKER@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh

Exact reviewed head: 6129@e4658d8e1cf5eff07b57fc8ca7eee75df319e0e2; immutable baseline/merge base: ecffa78cc68d1bb64724fd7c662eb89d04c6f1ad.

动机

维护绑定 Codex App 自动化线程的操作者,需要判断该线程最近是否按预期产生了活动。

此前状态只有最近线程活动,缺少与已安装定时配置的对照;本 PR 增加时间窗口,意图让操作者区分近期活动、超时和无可用证据。当前实现却会把另一线程或独立定时任务的配置当作本线程的预期。

真实临时 Codex 数据库和活动记录通过生产读取器证明:匹配配置可得到活动窗口,暂停配置保持未知,观察不会修改源文件;错误目标和非法间隔仍能产生误导性的 fresh。

这个只读活动诊断不启动、暂停或重试宿主,也不证明定时 prompt 被送达或任务已经完成。

绑定自动化的身份和间隔解析、语义清单需先修复;更大的真实 prompt 送达与 agent-start 证明仍由原宿主集成议题承担,状态 API 的新字段也不能代替可操作的产品诊断入口。

改动思路

保留现有 Python 宿主文件读取和只读投影是合理的,但配置必须绑定到实际线程,且 RRULE 不能另造一套较宽的规则;这两项修复应在本 PR 内完成。

本批只交付可信的绑定线程活动窗口,不宣称真实 prompt 送达;先修复身份、解析和清单,产品诊断入口与宿主送达资格继续保留明确边界。

入口仍是既有 Chat 状态路由:先读真实线程活动,再读安装配置,最后仅投影时间窗口。fresh/stale/missing/unknown 是诊断,不是新执行许可;线程身份必须来自 canonical binding,不能从相同 Goal/agent 文本替代。2倍时间容忍窗口仍未通过真实定时触发资格验证,不能据此宣布 issue3927 关闭。

具体改动

规格依据 https://github.com/loopx-project/loopx/issues/3927,固定 revision 2026-10-10T12:35:38Z。先读改动前规格;本 PR 中的文档扩写和作者完成声明不能自证验收。

  • Minimal reproduction:deferred;This PR is a read-only activity-window diagnostic, not a prompt receipt.
  • Expected behavior:not_met;Wrong target_thread_id/standalone cron and malformed RRULE still label unrelated manual activity fresh.

关键代码讲解

  • _automation_rrule_interval_minutes (loopx/codex_app_thread_activity.py:305):New substring/regex adapter supplies interval; invalid input can become one minute.
  • _installed_automation_intervals (loopx/codex_app_thread_activity.py:325):Reads ACTIVE automation files and joins only inferred Goal/agent, omitting kind/target.
  • build_host_delivery_window (loopx/control_plane/agents/host_thread_activity.py:256):Projects age into fresh/stale/missing/unknown without host effects.
  • ChatStatusRequestMixin._status (loopx/chat_status_api.py:126):Existing status caller now attaches diagnostic windows after actual activity observation.

还完整检查了 host window 的 missing/unknown reason、公开数据清理、原活动排序/上限、观察失败回退和221行新增测试。配置文件的ACTIVE筛选不能证明它属于实际线程;当前 sorted 后同 Goal/agent 覆盖也不应被当作唯一绑定证据。

正向路径:Canonical binding → real SQLite/rollout activity read → ACTIVE matching heartbeat interval → age projection → status readback。负向路径:Same Goal/agent, target different thread or kind cron → existing observer sees manual bound-thread activity → Goal/agent interval match → fresh diagnostic on unrelated source。

独立本地验证:实际源码测试 base38/head44 passed;同一真实宿主五例前后对照保留源文件字节和私有线程隔离。exact-diff canary 的5项直接检查通过,15项选中并执行,其中14项通过、1项失败,没有 skip。失败是上述两处 source I/O census,基线同一 full-tree semantic smoke 通过。

对主干的风险

  1. [P2] Bind the expectation to the actual heartbeat thread (loopx/codex_app_thread_activity.py:349)

触发:Same inferred Goal/agent in another heartbeat target or standalone cron while the bound thread has manual activity。证据:Status reports fresh although no relevant installed heartbeat was identified. The five-case real Codex SQLite/rollout/TOML fixture reproduces both wrong-target and standalone-cron cases. 最小修复:Join the canonical bound thread to heartbeat kind and target_thread_id; ambiguous unrelated matches must remain unknown and private IDs must stay internal. 回归:Matching vs wrong target/kind, multiple ambiguous files, paused and valid readback.

  1. [P2] Reject malformed RRULEs instead of inventing a one-minute interval (loopx/codex_app_thread_activity.py:317)

触发:FREQ=MINUTELY;INTERVAL=-3 or bad; malformed MINUTELYISH also matches。证据:Digit regex misses the provided invalid interval and falls back to1; negative fixture becomes fresh with expected_interval_minutes1, unlike canonical parser rejection. 最小修复:Reuse canonical recurrence parsing/contract; distinguish truly omitted INTERVAL from present-invalid, and require exact frequency. 回归:Negative/bad/zero interval, malformed frequency and duplicate fields versus canonical parser; valid cadence remains.

  1. [P2] Reconcile the source I/O manifest after shifting status imports (loopx/chat_status_api.py:19)

触发:The added imports shift two existing load_registry codec_read source locations。证据:Required semantic-vocabulary-drift smoke passes on base but fails on this exact head at both ChatStatusRequestMixin._status codec_read census rows. 最小修复:Regenerate and review the project registry I/O manifest for these two source entries, then rerun semantic smoke and exact-diff canary. 回归:Base/head same semantic-vocabulary-drift-smoke.py; head should pass with correct census, without relaxing scan or limits.

语义与 CI 对齐

开发时 advisory 先于 full-tree check 执行,新增 Enum 属于共享诊断词汇,不能因已注册/有类型就证明匹配语义正确。当前必需的本地 semantic census 失败与本次 imports 造成的行位置变化相符;应修复源清单并运行 uv run --extra test python examples/semantic-vocabulary-drift-smoke.py,以及 uv run --extra test loopx --format json canary premerge --from-git-diff --git-diff-base ecffa78cc68d1bb64724fd7c662eb89d04c6f1ad。不扩大阈值、不缩小扫描范围。

Actual Codex SQLite/rollout/TOML and production observer/projection. Status-route test has stubbed upstream collect_status and captured HTTP send; no actual scheduled App turn. 没有获取、轮询或等待 GitHub CI,没有对安装中的自动化、活动 Goal 或真实会话注入故障。配对观察使用固定时间和同一五例,基线没有新窗口;这种差异是有意新增诊断,三个错误输入的 fresh 则是当前回归。

我的整体评价

REQUEST_CHANGES。当前是有用的独立读回增量,long_horizon 的执行/重试/结算 owner 保留;user_experience 仍因已复现的健康或启动范围误述发生回归,故当前 exact head 不能批准。本批只交付可信的绑定线程活动窗口,不宣称真实 prompt 送达;先修复身份、解析和清单,产品诊断入口与宿主送达资格继续保留明确边界。

未来重构检查应在本批完成局部 recurrence 规则合并及 canonical binding,不需要新scheduler/receipt框架或无关语言迁移。发布的诊断入口仍是status API;更大产品诊断及真实送达receipt gap保留原issue,不以测试数量或“complete”标签关闭。 delivered recovery-review advice仅作为检查普通用户结果的线索,未继承其历史结论,也未声称记忆效用已验证。修复后重新读取head、复核完整PR和相关本地检查;平台审批、dismissal和merge另依其权限与readiness。

English verdict: REQUEST_CHANGES - 6129@e4658d8e1cf5eff07b57fc8ca7eee75df319e0e2. Bind the expectation to the actual heartbeat kind/target thread, reject malformed RRULEs without inventing an interval, and reconcile the source I/O census. Real SQLite/rollout/TOML negatives reproduce wrong-target/cron fresh and negative-interval one-minute fresh; baseline38/head44 tests pass, but exact-diff canary has one introduced semantic census failure among15 checks.

Comment thread loopx/codex_app_thread_activity.py Outdated
Comment thread loopx/codex_app_thread_activity.py Outdated
Comment thread loopx/chat_status_api.py
Address the review on the delivery window. A lane was matched to an installed
automation by inferred Goal/agent text alone, so an automation bound to another
thread -- or one that is not a heartbeat at all -- could be read as this lane's
cadence, and the window then reported fresh from an unrelated source.

Place a lane by its canonical binding instead: Goal, agent and the bound
target_thread_id, the way resolve_codex_app_automation_rrule places one, and
require kind = "heartbeat". Two automations claiming one lane resolve to
neither. A lane the Goal binds to more than one thread has no single identity
and is offered without one, so its window stays unknown. The binding is joined
from the Goal's coordination.thread_agent_bindings, and the thread id still
never reaches the window payload.

Read INTERVAL through the canonical rrule contract rather than a second, looser
dialect of it, so an rrule this adapter cannot read reports no interval instead
of a guessed one minute. Regenerate the project registry I/O manifest for the
two load_registry sites the added status imports moved.

Related to loopx-project#3927. Does not close it and does not name a root cause.

Signed-off-by: GZY-SUPER-HACKER <162807803+GZY-SUPER-HACKER@users.noreply.github.com>
@GZY-SUPER-HACKER

Copy link
Copy Markdown
Contributor Author

Thanks — all three are correct, and I reproduced each one before changing anything. Pushed as
19b92f1008eb87869701f3b5237e0b13c84768f2 on the same branch; base is still ecffa78cc.

[P2] Bind the expectation to the actual heartbeat thread

Confirmed. _installed_automation_intervals keyed only on the inferred Goal/agent text, so an
automation installed for another bound thread or for another kind of the same Goal was read
as this lane's cadence. My own fixture (one lane's heartbeat at INTERVAL=5, a second heartbeat for
the same Goal/agent bound to a different thread at 60, and a kind = "cron" automation at 7)
collapsed to a single 7 before the change.

Fixed by identifying a lane the way loopx.upgrade.resolve_codex_app_automation_rrule already does —
Goal, agent and target_thread_id — and by requiring kind = "heartbeat", which
control_plane/heartbeat/automation_upgrade.py:173 already requires of both stores. Two automations
claiming one lane now resolve to neither.

The binding is joined in attach_host_delivery_windows from the Goal's own
coordination.thread_agent_bindings, so HostDeliveryScope carries the thread id but nothing new is
serialized: the window payload still names no thread, and the test asserts the projection does not
contain it. A lane whose Goal binds more than one thread to the same (agent, surface) has no single
identity and is offered without one, which resolves to unknown rather than a neighbour's cadence.

Regression coverage added: own-thread match, other-thread, other-kind, two automations on one lane,
unbound thread, and the paused/valid readback that already existed.

[P2] Reject malformed RRULEs instead of inventing a one-minute interval

Confirmed, and the old adapter also contradicted its own docstring ("reports no interval, never
one"): INTERVAL=-3 and INTERVAL=bad both returned 1, and FREQ=MINUTELYISH matched as a
substring.

The adapter now implements the canonical contract instead of a looser dialect of it — the same
normalization, the same first-= split with last-duplicate-wins, an exact FREQ=MINUTELY, and a
positive integer INTERVAL read with the canonical integer conversion. scheduler_rrule_interval_minutes
in control_plane/scheduler/state.py is not a pure Python twin: it is an effect-runtime call
(state.py:89 → _operation_result → effect_runtime_result("scheduler.state.evaluate")), so
reaching the canonical parser through it would give the status route the scheduler runtime dependency
this design exists to avoid. Mirroring the contract keeps the two in step without that dependency; if
the canonical parser changes, this adapter is wrong and should be re-read against it.

One judgment call to flag. The canonical parser reports no interval when INTERVAL is omitted
(state_store.ts:141-157: pythonInteger(parts.get("INTERVAL") ?? "") is null). I followed the
canonical over the RFC 5545 default of 1, so an omitted INTERVAL now yields unknown rather than a
one-minute window. LoopX writes an explicit INTERVAL (state_store.ts:129-132), so I expect this to
be unreachable in practice — say the word if you want the RFC default for that one case instead.

[P2] Reconcile the source I/O manifest after shifting status imports

Confirmed: the added imports moved ChatStatusRequestMixin._status::codec_read:load_registry#1/#2
from 152/179 to 158/185. Regenerated with scripts/generate_project_registry_io_manifest.py; the
diff is exactly those two line values, both still codec_read, and no scan range or threshold was
touched. semantic-vocabulary-drift-smoke.py now passes, and
scripts/generate_project_registry_io_manifest.py --check reports the manifest current at 290 sites.

Verification

  • tests/test_host_thread_activity.py and the neighbouring status/registry suites: 142 passed.
  • ruff check on the changed files: clean. (ruff check tests loopx/ still reports the two
    pre-existing errors in capabilities/content_ops/cli.py and capabilities/native_chat/codex_auth.py,
    unchanged on the merge base.)
  • examples/semantic-vocabulary-drift-smoke.py: ok.
  • canary premerge --from-git-diff --git-diff-base ecffa78cc… on this Windows/cp936 box selects 15
    and reports 8 failures. All 8 reproduce identically on the untouched merge base — they are the
    Windows effect-runtime and CLI-stdout failures, not this change — and
    semantic-vocabulary-drift-smoke.py is now in the passing set.

Still open (unchanged from the PR body)

No real Codex App automation has been installed and no real stall observed, so the 2× tolerance
remains a judgment rather than a qualified value, and this batch still does not claim that a scheduled
prompt was delivered or that an agent started.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; xhigh

动机

通过工作区状态查看已绑定 Codex App Agent 是否近期有活动的用户和后续诊断消费者。
原来只能看到宿主活动,无法与该线程实际安装的定时周期一起判断陈旧程度。当前版本按 Goal、Agent 和绑定线程关联活动窗口,普通错误周期会显示未知;但一个超长数字周期即使属于另一个 Goal,也会让整个工作区状态请求返回 500。
正常绑定可得到基于实际安装周期的活动新鲜度,重复或错线程绑定保持未知;当前异常隔离缺陷会丢失整份状态,尚不能作为可靠诊断增量交付。
活动窗口不是任务正文已送达、Agent 已按协议启动或自动化健康的证明;本 PR 没有新增前端交互、宿主修复、权限或额度写入。
先让不可读周期安全降为未知且不影响其他 Goal;issue 3927 的正文摘要与首次 Agent-start receipt、真实宿主资格及对应用户界面仍由既有问题的后续边界负责。

改动思路

活动与安装清单解析留在现有 Codex provider,窗口是既有宿主活动的派生诊断,状态 API 只消费该投影;本地 RRULE reader 可避免引入调度执行依赖,但必须实现规范解析的失败语义。
本 PR 是已绑定 App lane 的只读诊断前置增量,不能关闭 prompt-delivery 问题;本轮要求在现有 provider 内有界处理整数转换失败,保留健康 lane 的完整状态,不扩大为宿主执行系统。
比较过不改、只在一个消费者补文案和复用已有 owner 的共享投影。不改会继续缺少一致状态;复制规则会让后续调整需要多点同步。现有 typed owner 与 provider 分工合理,派生字段由真实政策或宿主事件产生,不要求用户另写“准备好”的事实。诊断不能接管执行权,次要读取失败不能关闭不相关状态。做过相邻边界的有界重构判断:此处先修既有场景或 provider 的明确失败语义,不为未来框架扩大本 PR。

具体改动

规格来源 #3927 ,固定读到的版本 issue-3927-updated-2026-10-10T12:35:38Z 。prompt-delivery-and-agent-start 与 host-integration-diagnostic 均 deferred:活动时间加安装周期不足以证明正文送达,既有问题仍保留真实宿主 receipt 与诊断采用缺口。5 文件增加 865、删除 6 行,包括本地 provider RRULE 读取、派生窗口、status 挂接、focused tests 和 IO manifest。当前 head 把 join 改为 Goal/Agent/target-thread 且只接 heartbeat;重复自动化或同 Agent 多线程歧义保持 unknown。频率与正整数规范已修正,普通负数、缺 interval、MINUTELYISH 不再假定一分钟。I/O manifest 行范围已修正。完整 PR 与旧 review-to-head 分开核验;没有新增 frontend 展示,属于明确的诊断前置增量,不能据此关闭原 issue。

关键代码讲解

对主干的风险

[P2] 整数转换失败必须隔离在 provider(第 340 行)。 int(raw_interval) 在 safe-integer 检查之前,Python 3.11+ 的默认十进制位数上限会对 4301 位数字抛 ValueError。该调用也在 _installed_automation_intervals 的 TOML exception 边界之外,而状态请求扫描所有 ACTIVE heartbeat 清单。独立真实临时 SQLite/rollout/TOML 通过生产 _status 路径复现:本线程的坏清单和另一个 Goal/Agent/thread 的坏清单都使整份 status 返回 500;相同 base 工作负载返回 200,所有原输入字节不变。请保守界定或捕获转换失败,降为无 expectation/unknown,保持其他健康 lane 可读;不要全局调大 Python 限制。补 own/unrelated 超限、普通不可读与修正后恢复的路由回归。
15 行双侧 probe 的其余 13 行通过,覆盖多 Goal 使用不同周期、错线程、cron/PAUSED、重复、歧义、普通错误与修复后恢复。probe 使用 production handler/provider/observer 与真实源文件,collect_status 与响应捕获是 synthetic fixture;没有实际 Codex App 启动或修改活跃 Goal。比较 base 为 5de7109,whole diff/canary 的 merge-base 为 ecffa78,这三处生产 owner 在两 baseline 之间无差异。首次 focused 命令选了不存在路径,零测试退出,已保留;更正后 116 tests 通过,不将调用错误算作 PR 问题。全局 manifest 扫描的持续成本未测,不从单次诊断成功推出效率数字。
116 focused Python tests and all 15 selected/5 direct premerge checks passed; paired fifteen-case real SQLite/rollout/TOML probe still exposes two current HTTP 500 regressions. The first focused command named a nonexistent test path and ran no tests; the corrected command/result is retained separately.

语义与 CI 对齐

新增分类为既有 owner 的 vocabulary 扩展或本地派生诊断,开发时 advisory 在全树 semantic 检查之前运行;完整语义、可维护性及风险选定本地检查通过。它们不覆盖以上独立反例。没有查询或等待远端 CI。当前错误有精确触发、可观察结果、最小修复与回归路径,不能通过改预算、弱化断言或再加一份说明消除。保留各次失败与未测边界,未授权合并、安装或宿主控制操作不由评审结果提供。

我的整体评价

本轮结论为 REQUEST_CHANGES。诊断设计有正向潜力,但当前跨 Goal 的状态 500 对长程可靠性与体验均是 regression:一次不相关清单错误就失去全工作区读回,增加排错与恢复成本。先建立失败隔离,才能获得活动窗口减少盲查的效率收益。 本 PR 是已绑定 App lane 的只读诊断前置增量,不能关闭 prompt-delivery 问题;本轮要求在现有 provider 内有界处理整数转换失败,保留健康 lane 的完整状态,不扩大为宿主执行系统。 单次测试、截图和计数不能证明生产持续收益;原 broader acceptance 保持开放。重新评审需要当前 head 修复与上述决定性路径的验证。精确 head 6129@19b92f1008eb87869701f3b5237e0b13c84768f2。

English review

This is a useful read-only prerequisite, not completion of issue 3927: activity age plus an installed heartbeat interval cannot prove task-body delivery or agent start. The new head fixes prior wrong-thread/kind association, ordinary malformed recurrence parsing and IO-manifest drift. I independently ran 116 focused Python tests, all 15 selected/5 direct local canary checks and 15 paired real temporary SQLite/rollout/TOML cases. P2: int(raw_interval) at line 340 raises at Python’s default decimal digit limit before the safe-integer guard. Both a bound and an unrelated 4301-digit ACTIVE heartbeat manifest turn the whole status handler into HTTP 500; the same base cases return 200, and source bytes remain unchanged. Catch or conservatively bound conversion in the existing provider and degrade unreadable expectation to unknown without changing interpreter limits. Add own/unrelated malformed-manifest and correction cases through the status route. collect_status and response capture are synthetic; manifest IO, thread observation, projection and the production handler are real. No live App receipt, new frontend journey, long soak, or throughput claim is made. Current cross-goal failure is a negative effect on reliability and operator efficiency, despite the positive diagnostic idea.
Exact head: 6129@19b92f1008eb87869701f3b5237e0b13c84768f2.

English verdict: REQUEST_CHANGES

…refuses it

Address the second review on the delivery window. INTERVAL was converted with
int() ahead of the safe-integer guard, so a manifest whose interval runs past
Python's default decimal digit limit raised ValueError inside
_installed_automation_intervals. That call sits outside the per-manifest TOML
boundary, and the status route answers any local projection failure with one 500
for the whole workspace, so a single oversized manifest -- in this lane or in an
unrelated Goal -- cost every healthy lane its readback.

Rule the value out by its significant digits before converting it. Leading zeros
are stripped first because the interpreter counts them too, so the length check
now bounds exactly the digits int() is given. An out-of-range or unreadable
interval still reports no expectation, so its lane reads unknown.

Related to loopx-project#3927. Does not close it and does not name a root cause.

Signed-off-by: GZY-SUPER-HACKER <162807803+GZY-SUPER-HACKER@users.noreply.github.com>
@GZY-SUPER-HACKER

Copy link
Copy Markdown
Contributor Author

Thanks — I confirmed the three earlier findings landed, and this one is real. Fixed at
b6c51de241f182235776a653d0b99c41b652d263 on the same branch; base is still ecffa78cc.

[P2] Integer conversion failure must be isolated in the provider

Reproduced before changing anything. int(raw_interval) ran ahead of the safe-integer guard, so a
manifest whose INTERVAL exceeds Python's default decimal digit limit raised ValueError inside
_installed_automation_intervals. The call sits outside the per-manifest TOML boundary and the status
route answers a local projection failure with one 500 — I see exactly the response you describe:

the status route errored: LoopX status could not be projected for the workspace. {'status': 500}

Both variants reproduce: the lane's own manifest, and a manifest belonging to another
Goal/agent/thread, each cost every healthy lane its readback.

One case worth adding to the probe. The interpreter's digit limit counts leading zeros, so a
length bound applied to the raw field would still raise:

int("0" * 5000 + "5")  ->  ValueError: Exceeds the limit (4300 digits)

A value like INTERVAL=000…0005 is 5001 characters, is valid per the canonical regex, and converts to
5. I mention it because the obvious repair — checking len(interval_field) before the conversion —
passes the 4301-digit case and still fails on this one. The fix therefore strips the sign and the
leading zeros first, applies the length check to what remains, and converts only those significant
digits:

negative = raw_interval.startswith("-")
digits = raw_interval.lstrip("+-").lstrip("0")
if len(digits) > _MAX_SAFE_INTEGER_DIGITS:
    return None
interval = int(("-" if negative else "") + digits) if digits else 0

An out-of-range or unreadable interval still reports no expectation, so its lane reads unknown and
every other lane keeps its window.

Bound rather than catch, deliberately. A try/except ValueError around the conversion would also
work, but it would make correctness depend on an interpreter setting that can be changed elsewhere in
the process, and it would treat an unreadable value the same way as a legitimate one. Bounding the
digits mirrors what the canonical parser does — pythonInteger reports no value for anything outside
the safe-integer range — so the adapter keeps the canonical contract instead of gaining a second
failure mode. I did not raise sys.set_int_max_str_digits and did not change interpreter limits. The
route's broad except Exception is left as the backstop it is, not narrowed.

Verification

  • Unit: 9×16, 2**53, 4301 digits, 4301 digits negative, and 5000 leading zeros all report no
    interval; 2**53 - 1 and the leading-zero …0005 still resolve, so the bound did not swallow valid
    values.
  • Route, own / unrelated / corrected-after-fix: three regressions through the real _status handler,
    each asserting the route answers once with the workspace projection and never with an error.
  • Counter-experiment: reverting only the provider fix and keeping the new tests makes them fail with
    the 500 above — 6 failed / 2 passed. The two that pass are the 16-digit out-of-range values, which the
    existing guard already handled; only the digit-limit cases were broken. The fix is restored on the head
    above.
  • tests/test_host_thread_activity.py and the neighbouring status/registry suites: 150 passed.
  • ruff check on the changed files: clean. semantic-vocabulary-drift-smoke.py: ok.
    generate_project_registry_io_manifest.py --check: current at 290 sites.

Still open (unchanged)

No real Codex App automation has been installed and no live receipt observed, so the 2× tolerance
remains a judgment rather than a qualified value, the frontend journey is untouched, and this batch
still does not claim that a scheduled prompt was delivered or that an agent started. The broader
acceptance in issue 3927 stays open.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; xhigh

Exact head: b6c51de241f182235776a653d0b99c41b652d263. Immutable pre-change base: ecffa78cc68d1bb64724fd7c662eb89d04c6f1ad. Whole5-file diff (+1026/-6) and unchanged provider, scheduler and HTTP callers reviewed; verdict freshly derived.

动机

排查 Codex App 定时任务为何没有推进的操作者,需要分辨某个 Agent 最近是否在它自己的运行周期内有活动。

以前只能手工对比活动时间与配置,先前修订还会因一份超长周期配置让整份状态失败;现在相同状态入口能给出本 Agent 的时间窗,无法确认的周期返回未知,其他健康 Agent 的信息仍可读。

隔离真实 SQLite、活动记录和自动化配置的对照验证中,正确目标得到时间窗,错误目标与不合法周期显示未知,全部 17 个当前输入均返回 HTTP 200,源文件保持不变。

本批交付活动时间窗的只读排障信息;真实提示词送达凭证、自动唤醒、界面展示和整个宿主问题的验收不在本批。 活动新鲜不等于提示词正文已送达;真实宿主送达凭证和后续界面采用仍由现有问题 #3927 跟进。

改动思路

复用现有只读宿主活动和状态入口,时间窗是从当前文件派生的排障信息,不增加调度决策或独立健康服务。 当前边界是现有 HTTP 状态入口的可靠活动诊断;真实正文送达、界面采用和问题 #3927 的完整验收继续保留。

独立依据是 https://github.com/loopx-project/loopx/issues/3927,固定文本 revision issue-3927-updated-2026-10-10T12:35:38Z。prompt-delivery-and-agent-start 与 host-integration-diagnostic 两项完整验收均 deferred:本批只提供活动诊断,既有 #3927 继续负责真实正文送达和可观察的宿主凭证。此前 review5479227971 接受这个有用前置边界;不把未接受草案或作者完成声明当规格。

现有 TypeScript scheduler 保留周期决策 owner;Python 在已有宿主 IO/只读 projection 边界读 installed manifest。被动 status 不调用有副作用的 scheduler runtime。镜像只解释 canonical MINUTELY、安全正整数语义;不新增另一套执行策略。

具体改动

  • _automation_rrule_interval_minutes:精确 MINUTELY、首个等号和最后同名字段规则;有效位数先受限,再转数字。4301+位整数安全返回未知;大量前导零加5仍解析5,全零未知。
  • codex_delivery_expectations:使用 Goal、Agent、surface 与原绑定 thread 精确 join,仅 ACTIVE heartbeat 的可读唯一配置提供周期;其他线程、cron、暂停和重复来源不能冒充本 lane 的周期。
  • build_host_delivery_window:typed enum/dataclass 给出 fresh、stale、missing 或 unknown。fresh 表示观测活动在2倍周期内,不能证明提示词正文送达。
  • ChatStatusRequestMixin:现有 management status 先读活动,再附派生 window,thread ID 留在内部 join;没有重复读活动或写宿主。IO manifest 仅更新两处已有 site 的源代码坐标。全部五文件均已核验。

对主干的风险

当前191项相关 Python tests 全通过;相同 suites 在不可变 base160项通过。额外同一 harness 的 base/head 各17个真实状态入口场景:匹配目标、错误线程、cron/暂停、重复来源、负数/缺省/错误频率、超长数字、无关超长配置、缺失线程、两个 Goal 的5/90分钟周期、歧义绑定、修正后再读,以及大量前导零/全零。当前全部HTTP200、预期状态正确,全部源文件 hashes 未改变。既有有效目标正常,不是只验证更严格拒绝。

此前 exact19b92 的本人真实入口证据显示 own/unrelated 超长周期均导致workspace500;现在相同故障类保持200,未知周期不污染健康 lane。所有旧意见逐项映射:wrong-thread/kind、malformed周期、IO坐标以及超长int转换均有当前代码和测试覆盖。历史证据保留原 revision,没有冒充本轮执行。

当前5 direct +15 selected premerge checks通过,ruff与diff clean;开发期 advisory发现两个局部诊断enum,归属现有host projection,不是新共享执行 authority。最初调用了工具所在主工作树,分析 root 不对;已保留误调用记录,并用同一工具 API 显式指定当前工作树,随后重跑 full-tree semantic smoke通过。一次 pytest 文件名误选也已保留,由当前真实五个 suites 补齐;二者属于评审调用错误,没有算成 PR bug。当前 required checks无失败/skip;没有查询、轮询或等待GitHub CI。

真实边界是隔离SQLite/rollout/TOML与生产HTTP handler;upstream status输入和HTTP response capture是受控fixture。没有真实自动唤醒、正文送达、界面渲染或长期模型收益证据。重复读不增长持久状态,不改变配额、调度、绑定或结算;全目录扫描成本和2倍周期容忍度的长期效果尚未量化。

我的整体评价

APPROVE。有界效果/效率方向为正:沿现有入口减少人工对比和跨Agent误归因,修正一份坏配置使整个workspace不可读的故障,同时保留未知状态和原权限边界。长期运行的调度/承诺不受这个只读投影改变;完整宿主验收与长期净收益继续保留缺口。

未来相关重构检查:已有typed window/host provider归属合适;有限parser镜像有明确无副作用边界和回归覆盖,当前不扩成健康服务或RRULE框架。未新增执行authority,未实现UI,因此本批按独立可用的诊断前置阶段批准。旧blocking review的GitHub closeout、maintainer merge和安装分别核验。

English verdict: APPROVE — b6c51de241f182235776a653d0b99c41b652d263. This justified diagnostic prerequisite safely reports cadence-relative observed activity through the existing read-only status path; exact binding/heartbeat identity and unknown-source semantics are preserved.191 head/160 base tests,17 matched cases per revision,5 direct/15 selected premerge checks,ruff and corrected-root advisory followed by full semantic validation pass. Prior wrong-target/parser/IO and huge-integer findings are resolved at this head. Activity freshness does not authenticate prompt-body delivery; live host receipts, UI adoption and long-run utility remain with issue3927. Maintainer merge/installation authority is separate.

@huangruiteng
huangruiteng merged commit 34e35f9 into loopx-project:main Oct 11, 2026
23 of 30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants