The gap
A message can be delivered, read, and still never acted on — and nothing in the system records that an answer was owed.
Concrete case from today: an agent finished its work, asked its coordinator one gate question, and sat blocked for nearly three hours. The message was not unread. It was read and passed over, because it arrived looking like every other status update. No unread badge would have caught it, because it was not unread.
"Unread" and "unanswered" are different problems. We solve neither today: check_inbox is broken (#1471), and even when it works it answers "did you see it", not "do you owe someone a decision".
Asked for
1. Sender-declared severity on a message. The sender knows whether it is a status update or a blocking question. Something like severity: info | needs-answer | blocking, defaulting to info.
2. Boomerang — an unanswered needs-answer/blocking message returns to the recipient's attention after an interval, rather than decaying into scrollback.
3. The clearing rule, which is the load-bearing part:
A boomerang clears when the SENDER marks it answered — not when the recipient reads it, and not on a timer.
This is the whole design. Read-receipts would have marked today's case resolved while the sender stayed blocked. Only the blocked party knows whether they are unblocked. A recipient reply is evidence, not proof — a reply that does not answer the question should leave the flag standing.
4. An escalation ladder. If a blocking message is unanswered after N intervals, it should surface to the recipient's supervisor, not merely repeat at the recipient. A single unresponsive coordinator should not be able to strand a dozen agents silently, which is the failure mode this exists to prevent.
Design notes
Severity inflation is the obvious failure, and the mitigation is to bind severity to a stated cost. A sender should have to say what staying blocked costs — "this blocks a customer-facing answer" versus "this affects a doc comment" — so a recipient can order the queue by something real. Severity with no stated cost tends to converge on "everything is urgent".
Boomerang interval should scale with severity, not be global. A blocking question repeating every few minutes is noise; one that never repeats is the current behaviour.
The recipient needs a queryable open set, not just interruptions. "What am I currently blocking?" should be one call and should be the first thing a coordinator reads — today the equivalent is maintained by hand in a state file, and that has proven more current than any tool-based sweep.
Do not build this on read state. Read state is the wrong signal and is also currently unreliable (#1471).
Why this class of defect keeps recurring
This is the same shape as several other issues in this repo: a spawn that reports success and launches nothing; a heartbeat written by a timer rather than by progress; a status field reporting healthy over a dead queue; an empty result indistinguishable from a failed query. In each case a well-formed signal stood in for a fact nobody verified.
An unanswered question sitting in scrollback is that same shape, with the coordinator as the component that silently failed. The fix is the same in kind: make the state explicit and make its absence loud.
Acceptance
- A
blocking message unanswered past its interval visibly returns to the recipient, and does so again until cleared.
- Reading it does not clear it. Only the sender marking it answered clears it.
- A recipient can list everything currently blocked on them, ordered by declared cost.
- An unanswered
blocking message escalates beyond the original recipient rather than repeating indefinitely.
- Test the negative case explicitly: a message read but not answered must still boomerang. That is the case this issue exists for, and a test that only covers "delivered and answered" would pass today.
The gap
A message can be delivered, read, and still never acted on — and nothing in the system records that an answer was owed.
Concrete case from today: an agent finished its work, asked its coordinator one gate question, and sat blocked for nearly three hours. The message was not unread. It was read and passed over, because it arrived looking like every other status update. No unread badge would have caught it, because it was not unread.
"Unread" and "unanswered" are different problems. We solve neither today:
check_inboxis broken (#1471), and even when it works it answers "did you see it", not "do you owe someone a decision".Asked for
1. Sender-declared severity on a message. The sender knows whether it is a status update or a blocking question. Something like
severity: info | needs-answer | blocking, defaulting toinfo.2. Boomerang — an unanswered
needs-answer/blockingmessage returns to the recipient's attention after an interval, rather than decaying into scrollback.3. The clearing rule, which is the load-bearing part:
This is the whole design. Read-receipts would have marked today's case resolved while the sender stayed blocked. Only the blocked party knows whether they are unblocked. A recipient reply is evidence, not proof — a reply that does not answer the question should leave the flag standing.
4. An escalation ladder. If a
blockingmessage is unanswered after N intervals, it should surface to the recipient's supervisor, not merely repeat at the recipient. A single unresponsive coordinator should not be able to strand a dozen agents silently, which is the failure mode this exists to prevent.Design notes
Severity inflation is the obvious failure, and the mitigation is to bind severity to a stated cost. A sender should have to say what staying blocked costs — "this blocks a customer-facing answer" versus "this affects a doc comment" — so a recipient can order the queue by something real. Severity with no stated cost tends to converge on "everything is urgent".
Boomerang interval should scale with severity, not be global. A
blockingquestion repeating every few minutes is noise; one that never repeats is the current behaviour.The recipient needs a queryable open set, not just interruptions. "What am I currently blocking?" should be one call and should be the first thing a coordinator reads — today the equivalent is maintained by hand in a state file, and that has proven more current than any tool-based sweep.
Do not build this on read state. Read state is the wrong signal and is also currently unreliable (#1471).
Why this class of defect keeps recurring
This is the same shape as several other issues in this repo: a spawn that reports success and launches nothing; a heartbeat written by a timer rather than by progress; a status field reporting healthy over a dead queue; an empty result indistinguishable from a failed query. In each case a well-formed signal stood in for a fact nobody verified.
An unanswered question sitting in scrollback is that same shape, with the coordinator as the component that silently failed. The fix is the same in kind: make the state explicit and make its absence loud.
Acceptance
blockingmessage unanswered past its interval visibly returns to the recipient, and does so again until cleared.blockingmessage escalates beyond the original recipient rather than repeating indefinitely.