Skip to content

The silent-shard sweep imposes an undocumented liveness obligation on the engine #5

Description

@velo

Found while verifying the barrier fixes. Verdict-safe and self-healing — not a bug to fix now, but it must not be forgotten when the engine lands.

The obligation

Session.departSilentShards presumes dead any unreleased shard that holds no lease and has been silent for 3×5s. The 15s tolerance is calibrated to the barrier-poll cadence — but it is applied to shards in every phase, including phases where no cadence is mandated.

So the coordinator has silently imposed a contract on the engine: never be silent for more than 15s without holding a lease or being released.

Nothing documents that. The engine does not exist yet, so nothing honours it either.

The scenario, reproduced against a real container

Shard 0 waits at a barrier. Shard 1 is alive, mid-pass, in a ~20s gap between test classes — holding no lease — with a PENDING unit only it will claim.

At t+15s shard 1 is presumed dead. Shard 0 is answered DONE and released, irreversibly, while claimable work remains — the one property the release condition promises not to violate. Shard 1 then revives on its next claim, and runs and retries that unit entirely alone: zero rebalance. If shard 1 later dies for real, the work strands and the session reads INCOMPLETE.

Real triggers: a slow @AfterAll/@BeforeAll between classes, a long GC pause, a 15s network blip.

Every consequence is red-or-correct — never a false green — and a merely-partitioned shard self-heals on its next call. That is why it did not block the barrier PR.

Options when the engine is built

  • Raise the tolerance for shards not at a barrier. watermark == current pass is the observable "actually waiting" predicate. Care needed: the broad sweep also covers the post-restart mid-pass ghost, so narrowing it must not reopen that.
  • Document and honour the obligation engine-side — a keepalive, or claiming eagerly enough that a lease is always held across class boundaries.

Whichever is chosen, the engine and the sweep have to agree, and today only one of them knows the rule exists.

Smaller, related

LogRecord.departed does not record which kind of departure it was, so replayDeparted marks every replayed departure explicit. After a restart, an expiry- or presumption-departed shard that is actually alive comes back explicitly departed, so its barrier polls no longer revive it — only a claim will. Bounded, degraded-rebalance-only, never a wedge and never a wrong verdict. Recording the departure kind in the DEPARTED record would close it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions