Symptom
CI's tests job fails on :tramai-orchestration:test with timing-sensitive lifecycle assertions. Observed twice on unrelated exact heads in the same review cycle:
| Exact head |
Failing tests |
48e1dafe280b94b2688271d71d563b78f4b4448a (#407) |
ShutdownHookTest > recording observer captures lifecycle events during workflow execution and shutdown(), WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing() |
741a76fd96d9fd7fc74a9dc5d2d66400999795ad (#409) |
CheckpointPollerTest > poll failure is observable and cancellation is preserved() (java.lang.AssertionError at CheckpointPollerTest.kt:255), WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing() |
Two distinct failure modes tracked here
1. Orchestration timing assertions — test defect. CheckpointPollerTest,
WorkerShutdownCoordinatorTest, ShutdownHookTest fail on event-ordering / drain-completion
assertions with AssertionError. Evidence and attribution in the sections below. The fix
belongs in the tests themselves (deterministic waits instead of fixed windows).
2. Whole-suite watchdog exhaustion — CI calibration defect. The tests job background-runs
./gradlew test and polls it for 18 x sleep 40 = 12 minutes inside a timeout-minutes: 15
step, killing the build when it is still running. Ordinary successful runs of this suite reach
~12.8 minutes, so the containment threshold now overlaps normal runtime and the job is aborted
with no failing test and no tests completed tally. Observed twice on
a4eab8b628f8f037db37ec32f6373eea91baca21 (PR #412, a one-line YAML change), both times
Run tests exceeded 12 minutes — dumping threads and aborting with every downstream step
skipped. Being fixed separately as a calibration change (watchdog 18 -> 21 iterations = 14 min,
hard step timeout unchanged at 15).
They may share runner-load pressure as a contributing factor, but they are not the same
defect: one needs test-level determinism, the other needs CI budget calibration. Do not read
evidence for one as evidence for the other.
Evidence that this is a test/environment flake, not a regression
Hypothesis to investigate
Fixed timing windows (await/withTimeout budgets), scheduler contention under CI parallelism, and Gradle daemon load sensitivity on the runner. The drain assertions wait for a worker to finish, so a slower runner can observe a shorter event sequence than the fixture expects.
Constraint
Do not "stabilize" these opportunistically inside unrelated PRs — every such PR would otherwise carry an unaudited timing change. Handle it as its own change with a stated window/await policy.
Symptom
CI's
testsjob fails on:tramai-orchestration:testwith timing-sensitive lifecycle assertions. Observed twice on unrelated exact heads in the same review cycle:48e1dafe280b94b2688271d71d563b78f4b4448a(#407)ShutdownHookTest > recording observer captures lifecycle events during workflow execution and shutdown(),WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing()741a76fd96d9fd7fc74a9dc5d2d66400999795ad(#409)CheckpointPollerTest > poll failure is observable and cancellation is preserved()(java.lang.AssertionError at CheckpointPollerTest.kt:255),WorkerShutdownCoordinatorTest > shutdown drains a running execution to completion before finishing()Two distinct failure modes tracked here
1. Orchestration timing assertions — test defect.
CheckpointPollerTest,WorkerShutdownCoordinatorTest,ShutdownHookTestfail on event-ordering / drain-completionassertions with
AssertionError. Evidence and attribution in the sections below. The fixbelongs in the tests themselves (deterministic waits instead of fixed windows).
2. Whole-suite watchdog exhaustion — CI calibration defect. The
testsjob background-runs./gradlew testand polls it for18 x sleep 40= 12 minutes inside atimeout-minutes: 15step, killing the build when it is still running. Ordinary successful runs of this suite reach
~12.8 minutes, so the containment threshold now overlaps normal runtime and the job is aborted
with no failing test and no
tests completedtally. Observed twice ona4eab8b628f8f037db37ec32f6373eea91baca21(PR #412, a one-line YAML change), both timesRun tests exceeded 12 minutes — dumping threads and abortingwith every downstream stepskipped. Being fixed separately as a calibration change (watchdog 18 -> 21 iterations = 14 min,
hard step timeout unchanged at 15).
They may share runner-load pressure as a contributing factor, but they are not the same
defect: one needs test-level determinism, the other needs CI budget calibration. Do not read
evidence for one as evidence for the other.
Evidence that this is a test/environment flake, not a regression
LegacyPublicationDescriptions.kt)../gradlew :tramai-orchestration:test --tests "*CheckpointPollerTest*" --tests "*WorkerShutdownCoordinatorTest*"→ BUILD SUCCESSFUL on the same head.AssertionErroron event ordering / drain completion, not compilation or classpath errors.Hypothesis to investigate
Fixed timing windows (
await/withTimeoutbudgets), scheduler contention under CI parallelism, and Gradle daemon load sensitivity on the runner. The drain assertions wait for a worker to finish, so a slower runner can observe a shorter event sequence than the fixture expects.Constraint
Do not "stabilize" these opportunistically inside unrelated PRs — every such PR would otherwise carry an unaudited timing change. Handle it as its own change with a stated window/await policy.