Pear broker work must treat duplicate delivery as a normal failure mode. Renderer effects, dashboard reconnects, broker event-stream refreshes, and relay replay can all cause the same logical event to be observed more than once.
- Make lifecycle operations idempotent.
broker:start, agent registration, integration notifications, and mount/link setup should return or record whether they actually changed state before triggering side effects. - Coalesce concurrent starts or attaches with keyed in-flight promises. Repeated UI calls for the same project/root/channels should wait on the existing operation instead of starting another broker or event stream.
- Prefer stable event identity over content matching, and use the identity fields each event actually carries. Chat/broker events carry
event_id,id, orseq; dedupe those by identity AND content hash together.worker_streamPTY chunks do not carryseq/event_id— those are ephemeral and excluded from the daemon replay buffer, so any dedup keyed on them is structurally inert (it can never match). PTY chunks instead carry a cumulative per-worker byteoffset; dedupe them by(generation, offset)identity AND content hash, wheregenerationis the broker event-stream generation the delivering listener was attached under (it scopes the offset so a fresh worker stream's low offset can't collide with a stale remembered one). Never drop a PTY chunk on content alone (identical consecutive chunks are normal terminal traffic — repeated keystroke echoes, byte-identical TUI repaint frames) and never on identity alone (a repeated offset with different bytes is fresh output, not a replay — dropping it corrupts the escape-sequence stream into the stacked-duplicate-lines rendering corruption). When correlation metadata is absent (nooffset), deliver every chunk and log the blind spot loudly — do not silently claim a protection that isn't running. When in doubt, deliver: a rare duplicate repaints once; a dropped chunk mangles the screen until the next full repaint. - Scope live event listeners with a generation token when reconnecting or refreshing streams. Stale callbacks from an older listener must bail before publishing IPC events.
- Keep PTY stream delivery separately guarded in main and renderer code. Main should suppress duplicate
worker_streamchunks beforebroker:pty-chunk; renderer buffers should tolerate repeated chunk metadata as a final guardrail. - Do not post integration or launch metadata on reused broker sessions. Notify agents only after a real broker start, reconnect, or state transition, and make repeated payloads no-ops when possible.
- Add regression tests when touching broker start, event streaming, PTY buffering, spawned personas, or integration notifications. Include duplicate/replay cases, not just the happy path.
- Add low-noise telemetry for suppressed duplicates and a rate-limited loud warning (with a running count) when correlation metadata is missing, so both real replays and dedup blind spots are visible without flooding the terminal.
The broker daemon's PTY emulation (served as attach snapshots, observable via agent-relay-broker dump-pty) is the ground truth for what a worker terminal shows. The renderer's xterm grid can diverge from it (e.g. xterm reflow-scrolls on width resizes; PTY-side emulators don't), and diff-painting TUIs like Claude Code then preserve the divergence forever — they skip cells they believe unchanged, so stale glyphs bleed through new rows and stacked repaint frames accumulate. src/renderer/src/lib/terminal-reconciler.ts closes the loop: when a terminal is quiet it compares the viewport to the broker's plain snapshot and repaints from the self-framing ansi snapshot on confirmed divergence. Do not remove it after fixing any individual corruption vector — it is the convergence backstop for the whole class, and its gating invariants (quiet window, activity-serial recheck, dimension match, confirm-twice, rate limit) each guard a real re-corruption path documented in the module header. Confirmed divergence logs [terminal] viewport diverged from broker screen — that line firing is the signal a new creation vector exists and is worth hunting.
The gates themselves must not be able to disable the backstop silently or permanently:
-
Detection is reported, not just repair. The line above fires on CONFIRMED divergence, before the repair is attempted, and states the outcome (
repainted from snapshot,repair rate-limited,repair skipped: …) plus how many rows differ. This module is both the convergence backstop and the app's only always-on divergence detector, and those are different jobs: gating the only telemetry on a completed repair meant a divergence whose repair was deferred by the rate limit — or whose screen went busy again before the second confirmation — was healed-or-not in silence. It also made the backstop invisible to any observer sampling less than ~10s of quiet (quiet 1.5s + a check interval to align + a second confirming check), so a fully diverged screen could read back as zero telemetry, i.e. "clean".RECONCILE_DIVERGENCE_LOG_GAP_MSrate-limits the reports and MUST stay strictly belowRECONCILE_MIN_REPAIR_GAP_MS, or the repair-rate-limited case — the one worth hearing about — goes quiet again. -
Persistent dims mismatch escalates, never just skips. The dimension-match gate exists for resizes in flight, but a PTY that stays at a different size than the grid is the exact state that creates stacked-frame corruption (the TUI frames repaints for the wrong row count) — and it would also gate the reconciler off forever. After consecutive mismatched quiet checks the reconciler calls
onPersistentDimsMismatch, which invalidates the size-sync ack and re-asserts the rendered grid as the PTY size (logs[terminal] PTY RxC disagrees with grid RxC). -
Stranded predictions roll back instead of holding the quiet gate shut. Optimistic-echo predictions are only confirmed or dropped by server output; a keystroke the TUI swallows (or one typed into a hung agent) leaves
hasPredictionstrue with no echo ever coming, which would close the quiet gate permanently. AfterSTRANDED_PREDICTION_ROLLBACK_MSof output silence the registry rolls the predictions back (erasing only optimistic glyphs — confirmed bytes are untouched) and lets reconciliation proceed. -
A thrown check never kills the poll loop.
snapshotTerminaldegrades to null, but the IPC bridge itself can reject; the check loop catches and rate-limit-logs instead of dying on an unhandled rejection.