Skip to content

test-stability: orchestration timing assertions + tests-watchdog exhaustion (two distinct modes) #411

Description

@GionaGranchelli

Symptom

CI's tests job fails on :tramai-orchestration:test with timing-sensitive lifecycle assertions. Observed twice on unrelated exact heads in the same review cycle:

Exact head Failing tests
48e1dafe280b94b2688271d71d563b78f4b4448a (#407) ShutdownHookTest > recording observer captures lifecycle events during workflow execution and shutdown(), WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing()
741a76fd96d9fd7fc74a9dc5d2d66400999795ad (#409) CheckpointPollerTest > poll failure is observable and cancellation is preserved() (java.lang.AssertionError at CheckpointPollerTest.kt:255), WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing()

Two distinct failure modes tracked here

1. Orchestration timing assertions — test defect. CheckpointPollerTest,
WorkerShutdownCoordinatorTest, ShutdownHookTest fail on event-ordering / drain-completion
assertions with AssertionError. Evidence and attribution in the sections below. The fix
belongs in the tests themselves (deterministic waits instead of fixed windows).

2. Whole-suite watchdog exhaustion — CI calibration defect. The tests job background-runs
./gradlew test and polls it for 18 x sleep 40 = 12 minutes inside a timeout-minutes: 15
step, killing the build when it is still running. Ordinary successful runs of this suite reach
~12.8 minutes, so the containment threshold now overlaps normal runtime and the job is aborted
with no failing test and no tests completed tally. Observed twice on
a4eab8b628f8f037db37ec32f6373eea91baca21 (PR #412, a one-line YAML change), both times
Run tests exceeded 12 minutes — dumping threads and aborting with every downstream step
skipped. Being fixed separately as a calibration change (watchdog 18 -> 21 iterations = 14 min,
hard step timeout unchanged at 15).

They may share runner-load pressure as a contributing factor, but they are not the same
defect
: one needs test-level determinism, the other needs CI budget calibration. Do not read
evidence for one as evidence for the other.

Evidence that this is a test/environment flake, not a regression

Hypothesis to investigate

Fixed timing windows (await/withTimeout budgets), scheduler contention under CI parallelism, and Gradle daemon load sensitivity on the runner. The drain assertions wait for a worker to finish, so a slower runner can observe a shorter event sequence than the fixture expects.

Constraint

Do not "stabilize" these opportunistically inside unrelated PRs — every such PR would otherwise carry an unaudited timing change. Handle it as its own change with a stated window/await policy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions