Skip to content

fix(agent-status): count only agent work in stats, and read a Grok background subagent as working - #22474

Merged
brennanb2025 merged 7 commits into
mainfrom
brennanb2025/lead-status-pr-b2
Sep 24, 2026
Merged

brennanb2025 merged 7 commits into
mainfrom
brennanb2025/lead-status-pr-b2

Conversation

@brennanb2025

@brennanb2025 brennanb2025 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor
Files Added Deleted Net
Test 9 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​716 $\color{#cf222e}{\Huge{\mathbf{−}}}$​16 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​700
Prod 6 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​138 $\color{#cf222e}{\Huge{\mathbf{−}}}$​37 $\color{#1a7f37}{\Huge{\mathbf{+}}}$​101

ELI5

Since #22295 a status row can say "working" for two different reasons: the main agent is still on its turn, or the main agent has finished and one of its helpers (a subagent, or a background shell it left running) is still going. PR #22452 made the row carry the main agent's own state beside the combined one. This change goes through the four readers this slice owns, writes down which question each one is really asking, and moves only the ones that were asking the wrong one. Two readers keep reading the combined status because that is what they mean. One reader (the stats recorder) switches to "is an agent actually executing". The plugin event gains the new fact without changing anything it already said.

What Changed

Which question each reader asks, and what happened to it

Reader The question it means Verdict
Sidebar smart sort (resolveAttention) "How urgently should this worktree draw the eye" — its classes are named after what the row shows (Needs you / Done / Working) Keeps reading the combined state. Pinned by tests so a future mechanical migration is a deliberate change.
Activity unread badge (countActivityUnread) "How many rows in the Activity feed are unread" — it must mirror the feed, which builds its rows from the combined state Keeps reading the combined state. Pinned by tests.
Stats recorder (AgentSessionTransitionRecorder) "Was an agent executing" — it feeds "Agents spawned" and "Time agents worked" Migrated to "the row reads working and is not a watch loop". Only a settled main agent ever produces a watch loop, so this is exactly "the main agent's own turn runs, or a settled main agent's live subagent still holds the row".
Plugin event agent.status.changed A public event; state must keep its meaning Gains mainAgent as an optional field beside state.

Before and after, as the user experiences it

  • Stats pane, "Time agents worked". Before: after the main agent finished with a background shell still running (the row shows "Monitoring background tasks"), the clock kept counting for as long as the shell ran, so a dev server left up overnight counted as hours of agent time. After: the clock stops when the main agent's turn ends. Time a subagent spends running after the main agent finished still counts, as it did before: subagents are agents. That now includes Grok (next bullet). "Agents spawned" is unchanged when the main agent resumes after its subagents report back (that is not a new spawn). One case does change: a main agent that finishes into a background-shell window and later resumes now counts as a second spawn, where before the whole window read as one long turn.
  • Grok, a background subagent that outlives the main agent. Before: when Grok's main agent ended its turn with a background subagent still running (Grok runs subagents in the background by default), the sidebar showed "Monitoring background tasks", the same as for a leftover shell. After: the sidebar shows "Working", as it does for a Claude subagent that outlives its main agent, and that time counts toward "Time agents worked". A Grok background shell alone still shows "Monitoring background tasks" and does not count; with both a subagent and a shell running, the subagent wins. Completion notifications are unchanged: nothing is announced while the subagent runs, and one "finished" is announced when Grok's final turn ends with nothing outstanding.
  • Stats pane, a helper's approval or question wait (Codex, and Claude subagents). The clock pauses while the row waits on the user, whoever raised the prompt: a helper's wait stops it exactly like the main agent's own prompt, and the resume counts as another spawn, as it already does after the main agent's own prompt. This is unchanged from before for the row's combined status; the new main agent fact (which keeps reading "working" during a helper's wait) is deliberately not allowed to keep the clock running.
  • Stats pane, older remote hosts. An SSH host too old to send the main agent's own state already marked a finished main agent's background shell as a watch loop, so its dev-server-overnight window now stops the clock too, instead of counting as agent time.
  • Stats pane, after a restart. An agent that was still running when Orca restarted (a terminal agent with no main agent state, such as one only its terminal title reports) used to have its first live update after the restart dated from before the restart, so the whole downtime counted as agent time. It is now dated from when this run saw it.
  • Sidebar order and the Activity badge. No change. A worktree whose main agent has finished while a subagent still runs keeps sorting as Working, and keeps counting as one live working turn, because that is the row the user sees.
  • Plugins. A plugin subscribed to agent.status.changed now also receives mainAgent: { state, outcome?, stateStartedAt } when the row carries it. state is exactly what it was. mainAgent.stateStartedAt is stamped by the machine the agent runs on (a remote machine over SSH), unlike receivedAt.

The mechanism

  • One derivation, isAgentTimeAccruing, sits beside the status fold that produces the row: state === 'working' && workingMode !== 'monitoring'. The fold emits monitoring only when the main agent has finished and it was told only watch work still runs, and every lane goes through it (Codex never emits monitoring: it counts every child as agent work), so a working row that is not monitoring is exactly the main agent's turn or its live subagent work. Which children count as watch work is each lane's call, made through the shared child-work classifier: Claude and structured chat pass only shells and monitors, and Grok now maps each entry of its end-of-turn task list to the same kinds (a subagent entry is agent work, a shell entry is watch work; monitor entries stay excluded as before, since they can run indefinitely). A test runs every input the fold accepts and checks the derivation against the rule written in terms of the main agent and its children. A waiting or blocked row never accrues, whoever raised the prompt. It does not read mainAgent, so hosts too old to send it get the same answer. It is deliberately not a liveness answer: a background shell is still live work.
  • The recorder mirrors that answer per pane instead of the raw state, so a same-answer row is still a snapshot and never a transition, and the restored/replay gate is unchanged: a restored row never opens a session, whatever its mainAgent says. Each edge is dated only by clocks this machine stamped, chosen by the row alone: an edge that leaves working is a state change, so the row's own state clock dates it; an edge inside working (a watch loop starting or giving way to agent work, or the first live update after a restored row) leaves that clock on an older state start, so the evidence clock dates it. On a live update that just turned working, the two clocks are the same instant. The main agent's own stateStartedAt is never used for stats: over SSH it is the remote machine's clock.
  • The plugin projection is now one pure function with its own tests, and the payload schema admits the new optional object with the same bound the row normalizer applies, so a mainAgent the bus considers malformed cannot take the whole event down.

Restored rows, per reader

  • Recorder: a restored row with mainAgent.state: 'working' never opens a session (test, ablated).
  • Plugin event: a restored row projects to nothing, even when its mainAgent reads working (test, ablated).
  • Smart sort and unread badge: both read through the existing freshness gate, which refuses any restored row before it looks at a state. Pinned with tests that carry a restored mainAgent.state: 'working' row, and ablated by deleting that gate.

Why

  • Two readers stay on state on purpose. The smart sort's classes and the Activity feed's rows are presentation, and both are built from the combined status the row displays. Ranking a working row among the finished ones, or counting an unread "done" event the feed does not show, would make the badge and the order disagree with the rows. The alternative, migrating every consumer to mainAgent mechanically, was rejected for exactly that reason.
  • The recorder wants "is an agent executing", not "is the main agent's turn running". "Time agents worked" is plural: a subagent running after the main agent finished is agent time and always counted. Reading only the main agent's turn would have dropped that time and counted the main agent's resume after its subagents as a second spawn. Reading only the combined state is what let a background shell count as agent time. The derivation answers only the stats question; lifecycle decisions (keeping the machine awake, reaping idle sessions) must keep treating a background shell as live, which this derivation does not.
  • A wait on the user is not agent time, whoever asked. "Time agents worked" is an aggregate of agent work, so it excludes time blocked on the user, unlike a per-turn wall-clock duration; a helper's wait is treated exactly like the main agent's own.
  • Why read the combined row rather than mainAgent. An earlier revision decided from mainAgent and fell back to state without it. That gave the same answer on every row a current host produces, but made the recorder depend on a second fact with its own clock, and each review round found another edge dated by the wrong clock. Reading state and workingMode, which every host already sends, removes that dependency, and older hosts follow the watch-loop rule instead of being exempt.
  • Why derive rather than publish a fourth field. "Does this row accrue agent time" is fully determined by two facts the row already carries. A stored copy could disagree with what it is made of.

Linked Issue

Follow-up to #22452 (merged) and #22295. No separate issue. A late Codex root Stop keeping an inferred cancellation, which this PR originally carried, landed in #22452, so the plugin event's mainAgent.outcome is already correct on that lane.

Visual Proof

Run in the app (hidden dev build of this branch, real Grok 1.0.41, isolated profile). The sidebar order and the Activity badge are unchanged by construction and pinned by tests.

Grok background subagent after the main agent's turn ended: the row reads Working (before this change it read "Monitoring background tasks", and the time did not count under the new stats rule)

Grok row reads Working while a background subagent runs

Grok background shell after the main agent's turn ended: still Monitoring background tasks

Grok row reads Monitoring background tasks while a background shell runs

Stats pane. After the subagent run: 2 spawned, 1m worked (exact store total 73.7s for a 75s subagent). During the shell's ~90s monitoring window the exact total stayed at 82.0s (only the main agent's own 8.3s turn was added). After Grok's completion turn: 86.1s. The card shows whole minutes, so it reads 1m throughout.

After the subagent During the shell window After the shell
Stats after subagent Stats during shell window Stats after shell
Subagent finished: row settles to Done Grok row Done after the subagent finished

Testing

At the final head, rebased on current main:

  • pnpm tc exit 0; pnpm exec oxlint on the 15 changed files exit 0; pnpm run check:code-quality:changed passed (0 new findings).

  • pnpm test over src/main/stats, the Grok status-store test, the Grok hook-listener and completion-notification tests, the fold, main-agent parity and child-work liveness tests, src/main/plugins, src/shared/plugins, smart sort and Activity unread: 75 files, 590 tests passed.

  • New tests drive the real status store: a row saved as working and restored an hour later opens its first live session dated at the live event, not the saved clock; a Grok background subagent keeps one stats session open across the main agent's end of turn, and a Grok shell closes it. The fold test checks the accrual rule against every input the shared fold accepts, including a child waiting on a human.

  • Deletion ablations, run when each behavior landed during review (the last one at this head), each file restored from a saved copy and hash-checked. Every one went red for the pinned reason:

    • evidence-clock edge branch deleted: 7/39 (3660000 vs 60000 ms on the restore test);
    • monitoring exclusion deleted: 7/39;
    • the waiting guard in the every-input check deleted: 1/15;
    • Grok subagents classified as watch work again: 6 red, including the store-plus-recorder test (session closed at the main agent's end of turn);
    • Grok subagent entries dropped: 6 red, including an early completion notification;
    • plugin projection mainAgent deleted, and the restored-row skip deleted: both red;
    • smart sort and unread pins with the reader switched to mainAgent: both red.
  • A real Grok 1.0.41 Stop hook payload was captured to confirm that a background subagent is listed as backgroundTasks: [{ type: "subagent", status: "running", ... }].

  • In the app: the three scenarios in Visual Proof (at the pre-rebase head ba772157ea, which carries the same PR changes).

  • Platforms: macOS only. No platform branches.

  • I manually tested these changes locally (hidden dev build, real Grok: see Visual Proof)

  • Automated tests added/updated, or explained why not below

AI Disclosure

Author: @BrennanKB5

Review

  • Deferred, named: the keep-awake lease and repo-maintenance gate (the other follow-up to feat(agent-status): publish the main agent's own state beside the combined row state #22452) ask a different question from the stats: whether any work, a background shell included, is still live. They should derive that from the same published facts, not reuse isAgentTimeAccruing, which excludes a monitoring row on purpose.
  • Grok subagents: fixed in this PR. Grok's end-of-turn event lists each in-flight background task with a type of shell, monitor or subagent; Orca used to treat shell and subagent alike as watch work. It now feeds them to the shared child-work classifier, so the Grok row follows the same "any live agent work wins" rule as Claude and structured chat. A Grok pane on an SSH host whose relay predates this change keeps the old reading (Monitoring, not counted) until that relay updates, because the host normalizes Grok's events. Hook events Grok sends from inside a subagent's own session are still dropped, as before, so a Grok subagent's approval prompt is not seen: the row keeps reading Working and that wait counts as agent time, unlike a Claude or Codex helper's wait.
  • Older hosts: the recorder no longer reads mainAgent, so an older host's watch-loop row stops the clock like a current host's. Hosts from before mainAgent produced monitoring only after the main agent finished (Claude's pane resolver and Grok's stop handling), so this applies the same rule, not a guess.

Agent skill upstream boundary

  • Not applicable, or this change follows docs/reference/agent-skill-sharing-upstream-boundary.md and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.

Notes

  • Wire: mainAgent on the plugin event is a new optional field (rule 1 of the wire compatibility reference). Nothing a paired client and host exchange changes; the recorder and the plugin tap both read the host's own store.
  • SSH: the recorder and the plugin tap read the host's own store, which holds relayed rows too; nothing here is local-only.
  • Folder workspaces: no workspace-type branch is touched.

Checklist

  • This PR is small and focused
  • I explained what changed and why (ELI5, the user-facing before/after, the mechanism, and why over the alternatives)
  • Before/after screenshots or videos attached for UI changes, or N/A with reason
  • Self-reviewed for correctness, security, and performance
  • Cross-platform, SSH/remote, and path/shortcut impact considered (or N/A)
  • pnpm lint, pnpm typecheck, pnpm test, and pnpm build pass (or CI will cover; local preferred)

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important

One correctness gap in the new edge clock: a stop can be dated before the event that triggered it, dropping time from "Time agents worked" in a narrow Codex sequence. Details inline.

Reviewed changes

  • Shared execution derivation. isAgentExecutionOwed answers "was an agent executing" from state, workingMode, and lead, with a deliberate old-host fallback to state === 'working'.
  • Recorder reads executing. The per-pane mirror stores the derived boolean instead of the combined state, and agentExecutionEdgeAt chooses the clock for each edge.
  • Plugin event gains lead. A new pure projection adds the lead fact beside state; the payload schema admits it as an optional field and restored rows still project to nothing.
  • Renderer readers pinned, not migrated. Smart sort and the unread badge keep reading the combined state, with tests that record the choice.
  • Codex verdict carry-forward. A Stop keeps the inferred-interrupt verdict on the Codex root record; any other root event drops it, on both the hook and relay paths.

ℹ️ Nitpicks

  • docs/reference/agent-status-store.md is the canonical description of the store and its readers, but it says nothing about isAgentExecutionOwed or the recorder moving off the combined state. A sentence in "The lead fact" would keep the doc honest about which question each reader asks.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

Comment thread src/main/stats/agent-session-transition-recorder.ts Outdated

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ℹ️ No new issues in this delta — the commit is tests-only and the added tests are sound. The edge-clock thread from the previous review remains open: the new recorder test covers the approval-wait ordering, not the lead-settles-while-the-row-is-already-non-working ordering that thread describes, and agentExecutionEdgeAt is unchanged by this commit.

Reviewed changes

  • Added a recorder test pinning that a Codex child's approval wait (combined waiting) does not split a span while the lead's own turn runs. Verified meaningful by ablating the lead read in isAgentExecutionOwed: the test then reports 2 spawns / 45s instead of 1 / 60s.
  • Added an isAgentExecutionOwed assertion that a waiting row with a working lead still owes execution.

Pullfrog  | Fix it ➔ | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@brennanb2025
brennanb2025 force-pushed the brennanb2025/lead-status-pr-b2 branch from 9fdfbfb to 093ca10 Compare September 23, 2026 17:41
@brennanb2025
brennanb2025 force-pushed the brennanb2025/lead-status-pr-a branch from 3798c11 to 2b24918 Compare September 23, 2026 18:02
@brennanb2025
brennanb2025 force-pushed the brennanb2025/lead-status-pr-b2 branch 2 times, most recently from 093ca10 to 3901004 Compare September 23, 2026 18:07
@brennanb2025
brennanb2025 force-pushed the brennanb2025/lead-status-pr-b2 branch from 3901004 to 94739e4 Compare September 24, 2026 01:06
@brennanb2025
brennanb2025 changed the base branch from brennanb2025/lead-status-pr-a to main September 24, 2026 01:06
@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 13791143-5905-4fa0-af2f-daeaabdec9de

📥 Commits

Reviewing files that changed from the base of the PR and between ba77215 and 91439ef.

📒 Files selected for processing (3)
  • src/shared/agent-lead-status-fold.test.ts
  • src/shared/agent-lead-status-fold.ts
  • src/shared/main-agent-status-parity.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The change adds a shared predicate that determines whether an agent status accrues time. Grok background subagents now produce a working row, while shell tasks and active stop hooks produce monitoring status. Session transition recording uses the predicate and selects execution-edge timestamps. The plugin status event schema and projection support an optional main-agent fact. The main process emits the projected event only when a payload is produced. Tests cover these changes, activity unread counts, and sidebar attention.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to 91439

The change is mergeable with a narrow test-coverage follow-up: the OSC-repaint test does not protect its pinned-row-clock case. No current failure was established.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 16 functions across 15 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main changes: stats count only agent work, and Grok background subagents remain in the working state.
Description check ✅ Passed The description is comprehensive and covers the required change summary, rationale, linked issues, visual proof, testing, AI disclosure, compatibility notes, and checklist. The Linked Issue section re…
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 5c20594d-000d-466d-9d2f-5e5c1fd0173b

📥 Commits

Reviewing files that changed from the base of the PR and between 8352752 and 7798410.

📒 Files selected for processing (10)
  • src/main/plugins/plugin-agent-status-event.test.ts
  • src/main/plugins/plugin-agent-status-event.ts
  • src/main/startup/main-process-plugins.ts
  • src/main/stats/agent-session-transition-recorder.test.ts
  • src/main/stats/agent-session-transition-recorder.ts
  • src/renderer/src/components/activity/useActivityUnreadCount.test.ts
  • src/renderer/src/components/sidebar/smart-attention.test.ts
  • src/shared/agent-lead-status-fold.test.ts
  • src/shared/agent-lead-status-fold.ts
  • src/shared/plugins/plugin-events.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 2 remain after this review.

const local = new AgentSessionTransitionRecorder(osc)
local.onStatus(mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T))
local.onStatus(hook('done', T + 10_000))
local.onStatus(mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T + 12_000))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make the OSC-repaint reopen test use the unchanged row clock.

The test comment says the next hook row "restores the fact with its unchanged clock, which predates that close". Line 394 passes stateStartedAt: T + 12_000 instead. hook sets receivedAt to the same value. Both clocks equal T+12000, so the reopen edge is T+12000 whichever clock agentExecutionEdgeAt picks. A regression back to event.stateStartedAt would still produce 18000, so the test cannot catch it. To test the stated case, keep the row clock at T and put the observation time in receivedAt.

💚 Proposed fix
-    local.onStatus(mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T + 12_000))
+    local.onStatus(
+      mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T, {
+        receivedAt: T + 12_000
+      })
+    )
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
local.onStatus(mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T + 12_000))
local.onStatus(
mainAgentHook({ state: 'working', mainAgent: MAIN_AGENT_WORKING }, T, {
receivedAt: T + 12_000
})
)

Comment on lines +101 to +107
export function agentExecutionEdgeAt(event: AgentSessionStatusEvent): number {
const { mainAgent, state } = event.payload
if (!mainAgent || (mainAgent.state !== 'working' && state !== 'working')) {
return event.stateStartedAt
}
return event.evidenceObservedAt ?? event.receivedAt
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

The stop edge is still dated too early when the main agent settles under a row that is not working.

The issue from the earlier review is still present in the new code. Sequence: (waiting, mainAgent working) at T+1000, then the root Stops while a child still waits, which produces (waiting, mainAgent done). At that point isAgentTimeAccruing becomes false, so the recorder emits stop. agentExecutionEdgeAt then takes the first branch, because mainAgent.state is 'done' and state is 'waiting'. It returns event.stateStartedAt, which is the child's prompt time (T+1000), not the Stop time. The main agent's time between those two points is lost from totalAgentTimeMs.

This PR does not use the remote mainAgent.stateStartedAt. When mainAgent is present, date the edge with the local evidence clock, and keep the later of the two clocks:

🐛 Proposed fix
 export function agentExecutionEdgeAt(event: AgentSessionStatusEvent): number {
   const { mainAgent, state } = event.payload
-  if (!mainAgent || (mainAgent.state !== 'working' && state !== 'working')) {
+  if (!mainAgent) {
     return event.stateStartedAt
   }
-  return event.evidenceObservedAt ?? event.receivedAt
+  const observedAt = event.evidenceObservedAt ?? event.receivedAt
+  if (mainAgent.state !== 'working' && state !== 'working') {
+    // The row clock does not move when the main agent settles under an unchanged `waiting` row.
+    return Math.max(event.stateStartedAt, observedAt)
+  }
+  return observedAt
 }

Add a recorder test with this sequence: (working, main working)@T``, then (waiting, main working)@T`+1000`, then `(waiting, main done)` with `stateStartedAt: T+1000` and `receivedAt: T+5000`. Expect `totalAgentTimeMs` to be 5000.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
export function agentExecutionEdgeAt(event: AgentSessionStatusEvent): number {
const { mainAgent, state } = event.payload
if (!mainAgent || (mainAgent.state !== 'working' && state !== 'working')) {
return event.stateStartedAt
}
return event.evidenceObservedAt ?? event.receivedAt
}
export function agentExecutionEdgeAt(event: AgentSessionStatusEvent): number {
const { mainAgent, state } = event.payload
if (!mainAgent) {
return event.stateStartedAt
}
const observedAt = event.evidenceObservedAt ?? event.receivedAt
if (mainAgent.state !== 'working' && state !== 'working') {
// The row clock does not move when the main agent settles under an unchanged `waiting` row.
return Math.max(event.stateStartedAt, observedAt)
}
return observedAt
}

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important

The agentExecutionEdgeAt rewrite fixes the remote-clock skew and the pinned-row-clock reopen, but the stale-stop case from the previous review is unchanged: a stop is still dated by the row's clock when the row was already paused before the main agent settled. Re-verified against this head (1000 ms recorded where 5000 ms elapsed).

Reviewed changes

  • Rewrote agentExecutionEdgeAt to drop mainAgent.stateStartedAt (an SSH host's own clock) and date every edge from stateStartedAt / evidenceObservedAt / receivedAt.
  • Renamed isAgentExecutionOwed to isAgentTimeAccruing and reworded its doc to scope it to the stats question.
  • Added recorder tests for SSH clock skew, an OSC-repaint reopen, and a subagent reopening a monitoring row dated by the evidence clock.
  • Documented mainAgent.stateStartedAt as the host's own clock in the plugin payload schema.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

Comment thread src/main/stats/agent-session-transition-recorder.ts Outdated

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Reversed the helper-wait semantics: isAgentTimeAccruing now returns false for any non-working row before consulting mainAgent, so a child's approval or question pauses "Time agents worked" and its resume is a second spawn.
  • Simplified agentExecutionEdgeAt to !mainAgent || state !== 'working' → row clock, otherwise evidence clock; under the new fold the old mainAgent.state === 'working' evidence branch is unreachable, so the two changes stay consistent.
  • Replaced the recorder test that pinned one span across a child's approval wait with a parameterized test that pins the pause for both the main agent's own prompt and a child's, dated by the row clock.
  • Rescoped the derivation docs and added isAgentTimeAccruing cases for the pause, the watch loop, and the old-host fallback.
  • Verified the deltas: the five changed test files pass (100 tests), and the earlier stale-stop ordering can no longer strand an open session because executing is true only while the row is working.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Derivation now reads the combined row. isAgentTimeAccruing is state === 'working' && workingMode !== 'monitoring'; the mainAgent dependency and the old-host fallback are gone, so every host is read by the same rule.
  • Edge clock simplified to a state branch. agentExecutionEdgeAt dates an edge inside working by the evidence clock and an edge that leaves working by the row's own state clock; mainAgent.stateStartedAt is no longer read, which structurally retires the SSH clock-skew concern.
  • Recorder mirrors executing. classifyAgentSessionTransition keys on the derived boolean, and the recorder dates starts/stops only from this host's clocks.
  • Tests reworked around the combined row. The main-agent-fact suite was replaced by an old-host monitoring row that now stops the clock, a restored-row first-live start dated by the evidence, a real status-store hydration start, and the pause/SSH/reopen cases.
  • The load-bearing invariant is pinned. The fold test cross-products every foldAgentLeadStatus input and asserts the derivation matches "main agent turn, or settled main agent plus live agent child work", which is what licenses reading working + workingMode alone.

The load-bearing claim — monitoring implies a settled main agent — holds across the codebase: the fold only emits it under leadState === 'done', the Claude roster, Grok, and the structured lane all route through that fold, Codex never emits workingMode, and the payload normalizer admits monitoring only while state === 'working'.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

…they mean

Since #22295 a row's combined `state` reads `working` both while the lead's
turn runs and while a subagent or background shell outlives a settled lead.
The lead's own state now rides beside it (`lead`); each consumer in this slice
reads the question it actually asks.

- Smart sort and the Activity unread badge keep reading the combined state:
  their classes and rows are what the sidebar shows. Pinned with tests,
  including a restored `lead.state: 'working'` row that must never read live.
- The stats recorder asks "was an agent executing" and now reads a shared
  derivation (`isAgentExecutionOwed`): the lead's turn, or live agent child
  work holding a settled lead's row open. A settled lead's background shell
  no longer accrues "Time agents worked". Old hosts without `lead` fall back
  to today's read; restored and replayed rows still never open a session.
- The `agent.status.changed` plugin event gains `lead` as an optional field
  through one tested projection; `state` keeps its meaning and restored rows
  still project to nothing.
- A Codex root Stop that follows an inferred interrupt keeps the
  `cancellation` verdict, as the Claude lane already does at its turn
  boundary, on both the hook and relay paths.
… the recorded span

The recorder's move to the lead fact quietly changed one more story: a Codex
child's PermissionRequest turns the combined row waiting while the root's own
turn keeps running. The old state read closed the span there and minted a
second spawn on resume; the new read keeps one span, because the lead never
stopped. Pin it at both boundaries (the shared derivation and the recorder)
so the change is deliberate, not incidental.
…he accrual predicate to stats

The recorder dated a start by the producer's mainAgent.stateStartedAt. An SSH
host stamps that with its own clock while every stop is stamped locally, so each
span gained or lost the clock skew. The same clock also survives a row that
briefly lost the fact (an OSC repaint to another state), dating the reopen
before the close already sent, and a subagent reopening a monitoring row took
the row clock the hook lane pins to the main agent's turn start, re-billing the
whole watch-loop window. Edges now use the row clock when the row settles or
pauses and the evidence clock otherwise.

Rename isAgentExecutionOwed to isAgentTimeAccruing and state that it is the
stats question, not a liveness gate: it excludes watch loops, which lifecycle
gates must keep treating as live. Note on the plugin schema that
mainAgent.stateStartedAt is the execution host's clock.
…whoever asked

Time agents worked now accrues only while the combined row reads working. A
child's approval or question wait pauses the clock exactly like the main
agent's own prompt, and the pause edge is dated by the row's own clock.
…e edges by the row's own state

Time agents worked now accrues while the combined row reads working and is not
a watch loop. The shared fold emits monitoring only for a settled main agent, and
hosts that predate the main agent fact did the same, so this is the same answer
on every new-host row without reading mainAgent, and it applies the watch-loop
rule to older hosts too instead of billing their monitoring windows.

An edge that leaves working is dated by the row's state clock; an edge inside
working is dated by the evidence clock. This also stops a live repeat of a
hydrated working row (any row without the main agent fact, such as an OSC row)
from dating its start at the persisted state clock from the earlier runtime.
Grok's end-of-turn Stop lists each in-flight background task with its type
(shell, monitor or subagent). Orca filed a running subagent with the shells,
so a Grok subagent that outlived the main agent read "Monitoring background
tasks" and, with the stats recorder now skipping watch loops, stopped the
"Time agents worked" clock. Map shell and subagent entries to the shared
child-work kinds and let the shared liveness classifier decide: any live
subagent keeps the pane working, a shell alone or an active stop hook stays
monitoring, monitors stay excluded.
…ccrual check

Since the shared fold learned a child's human wait, a waiting child makes the row wait, so it
must not accrue agent time whatever the main agent is doing. The exhaustive check now includes
that input.
@brennanb2025
brennanb2025 force-pushed the brennanb2025/lead-status-pr-b2 branch from ba77215 to 91439ef Compare September 24, 2026 06:31

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Grok background tasks are now classified by kind. grokFiniteTaskKind maps a backgroundTasks[] entry's type (subagent → agent work, shell → watch work, everything else excluded), and grokChildWorkLivenessAfterStop aggregates them through the shared agentChildWorkLiveness, falling back to monitoring only for an active stop hook.
  • A live Grok subagent holds the row as plain working, not monitoring. The parity story and the listener tests pin a done main agent beside a listed subagent to { state: 'working' } with no workingMode, so its time counts toward "Time agents worked" and the sidebar reads "Working". A shell or an active stop hook still reads monitoring.
  • A discriminating end-to-end test. server-grok-background-status.test.ts drives the real server and recorder: a subagent keeps one open session across the main agent's end of turn, while a shell closes at the stop and opens a second session on the wake-up turn — the assertion would fail under the old "all finite tasks are watch work" rule.
  • Completion notifications unchanged. A new notification test pins silence while a Grok subagent outlives the main agent, then one announcement at the terminal turn; a cancel or session boundary still settles the pane.

I re-ran the four touched test files (64 tests) against this head — all pass. The "in-flight only" premise the classification rests on is supported by the captured fixture, which empties backgroundTasks once the task completes, so ignoring status is safe.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • Every-input accrual check now covers a waiting child. The fold cross-product test adds 'waiting' to the liveness inputs and asserts isAgentTimeAccruing is false whenever a child waits on a human, whatever the main agent is doing — correct for every fold branch.
  • Integration with the rebased base. The branch now sits on a main whose fold ranks a child's human wait above live work (childWorkLiveness === 'waiting' → state: 'waiting'). isAgentTimeAccruing already returns false for that row, so no code change was needed; the Grok path is unaffected because its candidates carry no state and so can never produce the new waiting arm.

I re-ran the six touched/dependent test files (96 tests) against this head — all pass.

Pullfrog  | View workflow run | Using DeepSeek Flash (free via Pullfrog for OSS) | 𝕏

@brennanb2025 brennanb2025 changed the title fix(agent-status): renderer and recorder consumers read the question they mean fix(agent-status): count only agent work in stats, and read a Grok background subagent as working Sep 24, 2026
@brennanb2025

Copy link
Copy Markdown
Contributor Author

Review status: ready for merge review

Head: 91439ef8e8, rebased on current main (retargeted from the merged #22452). CI: 22 checks pass, 13 path-filtered skips, none failing or pending.

What this PR now does, as the user sees it

  • Stats pane. "Time agents worked" counts the main agent's turn and live subagents after it finishes. It does not count a background shell or watch loop left running, and it pauses whenever the row is waiting on you, whoever raised the prompt. Every stats edge is dated by this machine's clocks.
  • Grok. A background subagent that outlives the main agent's turn now reads Working, and its time counts. A background shell still reads Monitoring background tasks.
  • Plugins. The agent.status.changed plugin event gains an optional mainAgent.
  • Sidebar order and Activity badge. Unchanged, and pinned by tests.

Review process

  • Loops. Six review loops, plus one architecture challenge after two loops in a row found the same class of bug (which clock dates a stats edge).
  • Architecture change. The recorder now reads only the row's own state and workingMode rather than the main agent's separate state and clock. That removed the bug class, and it also fixed a restart that could bill downtime.
  • Final checks. The final loop came back clean. The readiness checklist passed at both the start and the end, with no blockers.

Product decisions made during review

  • A helper's approval or question wait pauses "Time agents worked", the same as the main agent's own prompt. Resuming after the wait counts as another spawn, as it already does for the main agent's prompts.
  • Older hosts that don't publish the main agent's state follow the same rule: their watch loops stop counting too.
  • Grok's background subagents are classified as agent work. Grok runs subagents in the background by default.

Manual QA

  • Setup. Hidden dev build with an isolated profile and real Grok 1.0.41, run at pre-rebase head ba772157ea, which carries the same PR changes. All three scenarios passed; screenshots are in the PR body.
    • Subagent. A Grok background subagent showed Working, then Done.
    • Shell. A Grok background shell showed Monitoring.
    • Stats totals. Exact store totals: the 75s subagent added 73.7s. The ~90s shell monitoring window added nothing beyond the main agent's own 8.3s turn.

Known limits, deliberately out of scope

  • Grok's hook events from inside a subagent's own session are still dropped. So a Grok subagent's approval prompt is invisible, the row stays Working, and that wait counts as agent time. This is the same as on main.
  • A terminal status line (OSC 9999) that changes a Monitoring row's prompt drops its watch-loop mode. Nothing in the repo emits it.
  • When a row moves to a new pane key, the old key gets no clear event.

@brennanb2025
brennanb2025 merged commit 85642d0 into main Sep 24, 2026
35 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant