Found while verifying the barrier fixes. Verdict-safe and self-healing — not a bug to fix now, but it must not be forgotten when the engine lands.
The obligation
Session.departSilentShards presumes dead any unreleased shard that holds no lease and has been silent for 3×5s. The 15s tolerance is calibrated to the barrier-poll cadence — but it is applied to shards in every phase, including phases where no cadence is mandated.
So the coordinator has silently imposed a contract on the engine: never be silent for more than 15s without holding a lease or being released.
Nothing documents that. The engine does not exist yet, so nothing honours it either.
The scenario, reproduced against a real container
Shard 0 waits at a barrier. Shard 1 is alive, mid-pass, in a ~20s gap between test classes — holding no lease — with a PENDING unit only it will claim.
At t+15s shard 1 is presumed dead. Shard 0 is answered DONE and released, irreversibly, while claimable work remains — the one property the release condition promises not to violate. Shard 1 then revives on its next claim, and runs and retries that unit entirely alone: zero rebalance. If shard 1 later dies for real, the work strands and the session reads INCOMPLETE.
Real triggers: a slow @AfterAll/@BeforeAll between classes, a long GC pause, a 15s network blip.
Every consequence is red-or-correct — never a false green — and a merely-partitioned shard self-heals on its next call. That is why it did not block the barrier PR.
Options when the engine is built
- Raise the tolerance for shards not at a barrier.
watermark == current pass is the observable "actually waiting" predicate. Care needed: the broad sweep also covers the post-restart mid-pass ghost, so narrowing it must not reopen that.
- Document and honour the obligation engine-side — a keepalive, or claiming eagerly enough that a lease is always held across class boundaries.
Whichever is chosen, the engine and the sweep have to agree, and today only one of them knows the rule exists.
Smaller, related
LogRecord.departed does not record which kind of departure it was, so replayDeparted marks every replayed departure explicit. After a restart, an expiry- or presumption-departed shard that is actually alive comes back explicitly departed, so its barrier polls no longer revive it — only a claim will. Bounded, degraded-rebalance-only, never a wedge and never a wrong verdict. Recording the departure kind in the DEPARTED record would close it.
Found while verifying the barrier fixes. Verdict-safe and self-healing — not a bug to fix now, but it must not be forgotten when the engine lands.
The obligation
Session.departSilentShardspresumes dead any unreleased shard that holds no lease and has been silent for 3×5s. The 15s tolerance is calibrated to the barrier-poll cadence — but it is applied to shards in every phase, including phases where no cadence is mandated.So the coordinator has silently imposed a contract on the engine: never be silent for more than 15s without holding a lease or being released.
Nothing documents that. The engine does not exist yet, so nothing honours it either.
The scenario, reproduced against a real container
Shard 0 waits at a barrier. Shard 1 is alive, mid-pass, in a ~20s gap between test classes — holding no lease — with a PENDING unit only it will claim.
At t+15s shard 1 is presumed dead. Shard 0 is answered DONE and released, irreversibly, while claimable work remains — the one property the release condition promises not to violate. Shard 1 then revives on its next claim, and runs and retries that unit entirely alone: zero rebalance. If shard 1 later dies for real, the work strands and the session reads INCOMPLETE.
Real triggers: a slow
@AfterAll/@BeforeAllbetween classes, a long GC pause, a 15s network blip.Every consequence is red-or-correct — never a false green — and a merely-partitioned shard self-heals on its next call. That is why it did not block the barrier PR.
Options when the engine is built
watermark == current passis the observable "actually waiting" predicate. Care needed: the broad sweep also covers the post-restart mid-pass ghost, so narrowing it must not reopen that.Whichever is chosen, the engine and the sweep have to agree, and today only one of them knows the rule exists.
Smaller, related
LogRecord.departeddoes not record which kind of departure it was, soreplayDepartedmarks every replayed departure explicit. After a restart, an expiry- or presumption-departed shard that is actually alive comes back explicitly departed, so its barrier polls no longer revive it — only a claim will. Bounded, degraded-rebalance-only, never a wedge and never a wrong verdict. Recording the departure kind in the DEPARTED record would close it.