Skip to content

test: drop the remaining wall-clock deadlines from test waits - #26

Merged
MarcoDotIO merged 11 commits into
mainfrom
test/remaining-wall-clock-waits
Oct 1, 2026
Merged

MarcoDotIO merged 11 commits into
mainfrom
test/remaining-wall-clock-waits

Conversation

@MarcoDotIO

Copy link
Copy Markdown
Owner

Summary

#25 listed the test waits that still gave up at a wall-clock deadline. The macOS CI job runs about 3,150 tests in parallel on the xcode-27 runner, and the cooperative pool stalls for 5–8 s at a time, so any of these waits could fail even when its condition was about to hold. This PR converts the rest of that list using #25's pattern:

  • Poll while !Task.isCancelled with no deadline, using try? await Task.sleep.
  • On cancellation, Issue.record a labelled AsyncWaitTimeoutError and then throw it. Swift Testing drops errors thrown after a time-limit cancellation, so the recorded issue is what names the wait that hung.
  • Put .timeLimit(.minutes(1)) on every suite or test that reaches one of these waits.
  • Replace a timeout that only separated "short path taken" from "long default used" with a value past the time limit, so a regression shows up as a hang.

OpenClawKitTests

Wait Change
ProviderStreamingCancellationTests.waitUntil(seconds: 5) Uses the shared waitUntil; suite time limit. The request timeout (policy and URLSessionConfiguration) is now 1 h, so a cancellation that never reaches the request fails as a hang instead of passing when the old 60 s timeout fired.
ModelRouterStreamingFallbackTests 5 s loop Uses waitUntil; suite time limit. The two adjacent unlabelled polls in the same suite use it too.
RealtimeTalkRelaySession._test_waitForStartupCancelled(timeoutSeconds: 1) This was not a racy wait. stop() sets isClosed synchronously on the main actor, so the hook returned .cancelled before any timer was armed. The parameter is removed and the hook passes a 0 s timeout, so the test no longer depends on a wall clock at all. A session that is not closed now reports .failed immediately instead of after 1 s. The test has a per-test time limit.
OpenClawChatUITests (1 s) Done on #25 (45849eb), which this branch is on top of.

OpenClawLinuxRuntimeTests

  • New TestAsyncHelpers.swift, the same shape as the one in OpenClawKitTests. It replaces AgentLoopHardeningTests.waitUntil, which did 500 × 10 ms, recorded an unlabelled issue, then returned so the test carried on. Its three callers now pass labels; they already had per-test time limits.
  • GatewayServerTestHarness:
    • Recorder.waitFor (5 s) now takes a label and waits until the time limit.
    • collect had the same 5 s bug and 12 callers. It wasn't on the list, but it's in the same file, so it's converted here too. When the time limit cancels it, the for await ends and the helper records and throws.
    • .timeLimit(.minutes(1)) is added to the seven gateway suites that call these helpers.
  • One deliberate negative wait is kept. GatewayScopeAuthorizationTests checks that an unregistered connection gets nothing in 200 ms. It now uses frames(_:arrivingWithinMs:). The old call could never fail: whenever the 200 ms timer won the race, the task group returned [] and discarded the frames it had collected. The new helper returns what actually arrived.
  • AutomationHardeningTests (10 s × 2): both use waitUntil.
    • addingAJobWakesTheSleepingLoop used its 10 s deadline to tell "woken by the job change" apart from "slept out the 60 s idle cap". Without that deadline, the 60 s cap would race the 60 s time limit.
    • A new DEBUG-only hook, CronScheduler._test_setMaximumSleepSeconds, raises the cap to 1 h in that test. A missed wake now hangs until the time limit and fails.
  • SignInWithChatGPTSessionTests (10 s loopback-page wait): uses waitUntil; suite time limit.
  • MCPHardeningTests (3 s + 2 s, plus elapsed < 1.5 s, "no request-timeout stall"):
    • Both polls use waitUntil.
    • The request timeout goes from 500 ms to 1 h, and the elapsed assertion is removed. A call or refresh that stalls behind the reader now hangs until the time limit. MCP requests end on cancellation (withTaskCancellationHandler → cancelRequest), so the time limit does end them.
  • ChannelAdaptersLinuxSmokeTests.poll (15 s): removed in favour of waitUntil; suite time limit.

Timeouts deliberately kept

  • The 200 ms negative window above.
  • #expect(throws: .timedOut) / .timeout(...) cases, such as SIWC timeout: 1 and the MCP legacy SSE connect timeout. Those timeouts are the behaviour under test.

Stacked on #25 (which is stacked on #24). Merge those first; until then this diff also includes their commits. Only the last two commits are new: 7661599 (OpenClawKitTests + relay hook) and 3b3466e (Linux runtime).

Test plan

  • swift build --build-tests
  • Scripts/lint-swift.sh: 0 violations
  • Targeted swift test --filter over the 17 affected suites: 96 + 29 tests pass, in about 1 s each.
  • Pool starvation, using scratch tests that were not committed. Each run has one starver in the same test target, spawning activeProcessorCount * 2 (28) Task.detached(priority: Task.currentPriority) { usleep(6_000_000) } blockers, which gives about a 12 s stall.
    • Control: a copy of the old Date() < deadline loop with a 5 s deadline fails after 12.0 s. waitUntil on the same condition passes after 12.0 s.
    • Starver next to the real tests, all passing after 12.0–12.5 s (so they were stalled):
      • cancellingBeforeResponseHeadCancelsTheRequest
      • both ModelRouterStreamingFallbackTests cancellation tests
      • "closed relay does not wait for startup ready"
      • the two AgentLoopHardeningTests abort tests and duplicateRunIDsJoinTheActiveRun
      • all 6 GatewayEventStreamTests
      • the collect callers in the scope, approval, presence, organization, wire-shape and branch suites
      • smsWebhookDeliversSignedMessage
      • "Browser sign-in completes through the loopback listener"
      • toolsListChangedOnTheServerStreamRefreshesTheCatalog
    • The two automation tests finished before the starver took hold, so they were checked with scratch copies that spawn the blockers inline right after start() / addJob. The new versions pass after 12.1 s and 12.3 s. The old 10 s loop also passed in this run, because it re-reads the runs after each sleep before checking the deadline. It was converted anyway to remove the wall clock.
  • A wait that never holds ends at 60 s with Time limit was exceeded plus Timeout waiting for: <label>. Checked for the Linux waitUntil, collect and Recorder.waitFor.
  • Mutation checks, reverted afterwards:
    • CronScheduler.reschedule() with the wake removed: addingAJobWakesTheSleepingLoop fails at 60 s with Timeout waiting for: job run after the wake.
    • MCPClient awaiting the list_changed handlers on the reader, which is the deadlock the test guards against: toolsListChangedOnTheServerStreamRefreshesTheCatalog fails with Time limit was exceeded. With the old 500 ms request timeout, the same bug would have recovered after a request timeout and been caught only by the elapsed bound.
  • Full swift test on macOS:
    • OpenClawKitTests passed (3,150 tests in 341 suites).
    • OpenClawKitE2ETests passed (20 tests).
    • OpenClawLinuxRuntimeTests: 581 of 582 passed. The one failure was SpotlightMemoryTests.replyCoordinatorRoutesRepliesToOneInvocationAtATime, a known flake that predates this branch; this PR doesn't touch it. Re-running the target passed all 582 tests.

Not addressed here

A sweep for similar patterns found more waits, outside the list this PR worked from. They need a case-by-case audit, because some are yield loops or intentional bounds:

  • Iteration-count polls (for _ in 0..<N { … Task.sleep … }):
    • ApprovalQuestionBrokerTests
    • AgentRuntimeExtensionsTests
    • ExecApprovalHardeningTests
    • GatewayRunLifecycleTests
    • RuntimeIntegrationWiringTests
    • ToolsGatewayMethodsTests
    • GatewayServerChatUICompatTests
    • GatewayNodeSessionRouteTests
    • CoreAIModelRuntimeTests
    • VoiceNoteRecorderTests
    • SystemStateReportingTests
    • MediaPipelineTests
  • Elapsed-time assertions:
    • MCPHardeningTests.legacySSEConnectTimesOutWithoutAnEndpointEvent (< 2 s)
    • MCPStdioTransportTests (< 5 s, < 3 s)
    • SpotlightMemoryTests (< 3 s, < 2 s)
    • GatewayRunLifecycleTests (< 5 s)
    • AutomationHardeningTests cron search (< 1 s, synchronous)
    • ChannelAutoReplyTests
    • ModelRoutingTests
  • Product timeouts used as positive-path bounds:
    • SIWC signIn(timeout: 20)
    • A2A replyTimeoutMs: 5_000 in the channel smoke test
    • runtime.wait(timeoutMs: 10_000)

🤖 Generated with Claude Code

MarcoDotIO and others added 11 commits September 30, 2026 10:14
anotherAgentsRunBeforeTheAckNeverHijacksTheIntent and
preAckEventsOfTheAckedRunAreReplayed failed on main (CI run 36644330107,
xcode-27 runner) with "Caught error: CancellationError()" after ~8 s; the
same tree passed on PR #20. waitForSend polled the requester every 5 ms and
threw once a 5 s ContinuousClock deadline passed. With ~3,150 Swift Testing
tests saturating the cooperative pool, the host.send task (and the poller
itself) did not run for more than 5 s, so the deadline expired although
nothing in GatewayOpenClawIntentHost.send is slow. Blocking every cooperative
thread for 6 s reproduces the failure locally.

- OpenClawAppIntentsRunMatchingTests: HeldChatSendRequester yields each
  chat.send's params to an AsyncStream the moment the request arrives, and
  waitForSend awaits that stream. Assertions are unchanged.
- GatewayNetworkConnectionTransportTests: the loopback NWListener start waits
  on stateUpdateHandler (ready/failed/cancelled) instead of polling
  listener.state against the same 5 s deadline, which also had no final
  re-check after a late wake-up.
- WatchNodeClientTests: eventually() drops its 5 s deadline. Its conditions
  read production client state that has no change hook, so it still polls,
  but slowness no longer fails it.

Each suite gains .timeLimit(.minutes(1)) for hang protection. Every wait
ends on cancellation, so a time-limit overrun reports "Time limit was
exceeded" instead of hanging.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
keepalivePingIsBoundedWhenNoPongArrives and
handshakeTimeoutOptionBoundsTheWholeHandshake failed on PR #22's macOS job
(Xcode 27) with `ContinuousClock.now - start < .seconds(5)` at ~5.7 s. A
trivial test in the same window took 5.3 s, so the runner stalled; the
timeouts under test were 50 ms and 100 ms.

- keepalive ping: the socket never pongs, so the thrown URLError already
  proves the ping deadline ended the wait; a time limit catches a hang.
- handshake option: the fallback budget is raised to an hour through
  _test_setConnectTimeoutSeconds, so a channel that ignored the 100 ms option
  would trip the time limit instead of finishing at the 30 s default.
  Checked with a scratch test that drops the option: it fails with "Time
  limit was exceeded" and does not hang.

Both tests take .timeLimit(.minutes(1)) and pass while every cooperative
thread is blocked for several seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ot elapsed time

issuedTokenPersistenceNeverBlocksTheChannelActor asserted the actor answered
within 3 s while a token write waited on another connection's SQLite lock.
A saturated test pool can stall the run for longer than that, and a time
limit alone would miss the regression it guards (a write on the actor holds
it for SQLite's 30 s busy timeout, less than the one-minute limit).

The test now relies on ordering. It releases the lock only after the actor
answers, so the token can reach disk only if the actor answered while the
write was pending. A write on the actor fails with SQLITE_BUSY first, the
token never lands, and the final wait trips the new suite time limit.

The 200 ms sleep that let hello-ok reach the write is replaced by a DEBUG
hook, _test_setDeviceTokenPersistenceStartedHandler, that fires on the actor
just before the persistence hop. Once the hop starts, shutdown cannot stop it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… limit

The three tests asserted wall-clock bounds (2 s, 5 s, 1 s) that a saturated
test pool can overrun. Each now has a one-minute time limit instead, and
each makes sure a regression ends on the limit's cancellation instead of
hanging the run:

- file fetch refuses a FIFO: a blocking open(2) never returns. The
  cancellation handler opens the FIFO's write end once to release it.
- operationThatIgnoresCancellationStillTimesOut: the parked operation is a
  gate the test opens on cancellation (and afterwards), not a
  never-resumed continuation that a loser-joining race would wait on
  forever.
- total deadline can fire before the session starts: URLSession's own
  timeouts move past the limit, so only the fetcher's zero-second
  deadline can end the fetch. The fetch is awaited through a no-deadline
  AsyncTimeout race, because a deadline lost before start would leave it
  unresumed even on cancellation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aits

gatewayCoreWaitUntil gave up after 10 s (15 s at two call sites) and the
ChatViewModelSessionActionTests helpers after 15 s. A pool stall adds to
those waits, so they could time out even though the condition was about to
hold.

Both now wait until the condition holds or the test is cancelled. Every
suite that uses them has a one-minute time limit. gatewayCoreWaitUntil
records its GatewayCoreWaitTimeout as an issue before throwing, because
Swift Testing drops errors thrown after a time-limit cancellation and the
label says which wait hung. waitForForkStart awaits the gate's stream
directly instead of racing it against a sleep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
waitUntil gave up after 15 s (7, 10 or 30 s at five call sites). The macOS
CI job runs about 3,150 tests in parallel and the cooperative pool stalls
for 5 to 8 s at a time, so a wait could time out even though the condition
was about to hold.

waitUntil now polls until the condition holds or the test is cancelled, and
the timeoutSeconds and now: parameters are gone (nothing injected a clock,
and with no deadline there is nothing to measure). Every suite whose tests
reach waitUntil, directly or through a file-local helper, has a one-minute
time limit; ChatViewModelSessionActionTests already had one. On cancellation
the helper records AsyncWaitTimeoutError as an issue before throwing it,
because Swift Testing drops errors thrown after a time-limit cancellation
and the label says which wait hung.

No call site relied on the timeout: nothing expects AsyncWaitTimeoutError,
no wait runs in a child task that the test cancels, and the one try? wait is
cleanup in a catch block that rethrows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nnect tests

ChannelAdaptersE2ETests.waitFor gave up after 15 s and only recorded an
unlabelled expectation failure. reconnectFailureSchedulesAnotherAttempt
polled against a 10 s deadline. A pool stall can outlast either one.

waitFor now takes a label, polls until the condition holds or the test is
cancelled, and records a WaitTimeout issue before throwing it. The reconnect
loop polls until cancelled. The suite and the reconnect test each have a
one-minute time limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OpenClawChatUITests had its own waitUntil that gave up after 1 s and
returned false. On the macOS CI job for this PR, a pool stall of about
10 s ran past it: both view-model tests failed after about 12 s, and the
bootstrap expectations that followed failed with them.

The tests now call the shared deadline-free waitUntil with a label for
each wait, and the suite has a one-minute time limit. The private helper
is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ay tests

- ProviderStreamingCancellationTests polls with the shared deadline-free
  waitUntil under a one-minute suite time limit. Both request timeouts are
  now an hour, so a cancellation that never reaches the request fails as a
  hang instead of passing once the 60 s request timeout fires.
- ModelRouterStreamingFallbackTests waits for the stream termination (and
  the two stream/generate starts) with waitUntil, under a suite time limit.
- RealtimeTalkRelaySession._test_waitForStartupCancelled() drops its
  timeoutSeconds parameter. A closed session answers before any timer is
  armed; a zero timeout makes one that is not closed fail at once rather
  than after a 1 s wall-clock wait.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A shared TestAsyncHelpers.swift replaces AgentLoopHardeningTests'
  waitUntil (500 x 10 ms, recorded an unlabelled issue and returned). It
  polls until the test is cancelled, then records and throws
  AsyncWaitTimeoutError(label:).
- GatewayServerTestHarness.collect (5 s) and Recorder.waitFor (5 s) take a
  label and wait until the time limit. The one deliberate negative wait now
  uses frames(_:arrivingWithinMs:), which returns what arrived; the old
  helper returned [] whenever its timer won, so that assertion could never
  fail.
- ChannelAdaptersLinuxSmokeTests.poll (15 s), the SIWC loopback-page wait
  (10 s), the automation run waits (10 s) and the MCP list_changed waits
  (3 s + 2 s) use waitUntil.
- addingAJobWakesTheSleepingLoop pushes the scheduler's idle cap to an hour
  through the new DEBUG hook CronScheduler._test_setMaximumSleepSeconds, so
  a missed wake hangs instead of racing the 60 s cap against the time limit.
- The MCP list_changed test uses an hour request timeout and drops its
  "< 1.5 s" elapsed assertion: a stall behind another request now hangs.
- .timeLimit(.minutes(1)) is on the nine suites that reach these waits and
  had no limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MarcoDotIO
MarcoDotIO merged commit 7e21681 into main Oct 1, 2026
19 checks passed
@MarcoDotIO MarcoDotIO mentioned this pull request Oct 1, 2026
2 of 3 tasks
MarcoDotIO added a commit that referenced this pull request Oct 1, 2026
* test: wait for events instead of a 5 s wall-clock deadline

anotherAgentsRunBeforeTheAckNeverHijacksTheIntent and
preAckEventsOfTheAckedRunAreReplayed failed on main (CI run 36644330107,
xcode-27 runner) with "Caught error: CancellationError()" after ~8 s; the
same tree passed on PR #20. waitForSend polled the requester every 5 ms and
threw once a 5 s ContinuousClock deadline passed. With ~3,150 Swift Testing
tests saturating the cooperative pool, the host.send task (and the poller
itself) did not run for more than 5 s, so the deadline expired although
nothing in GatewayOpenClawIntentHost.send is slow. Blocking every cooperative
thread for 6 s reproduces the failure locally.

- OpenClawAppIntentsRunMatchingTests: HeldChatSendRequester yields each
  chat.send's params to an AsyncStream the moment the request arrives, and
  waitForSend awaits that stream. Assertions are unchanged.
- GatewayNetworkConnectionTransportTests: the loopback NWListener start waits
  on stateUpdateHandler (ready/failed/cancelled) instead of polling
  listener.state against the same 5 s deadline, which also had no final
  re-check after a late wake-up.
- WatchNodeClientTests: eventually() drops its 5 s deadline. Its conditions
  read production client state that has no change hook, so it still polls,
  but slowness no longer fails it.

Each suite gains .timeLimit(.minutes(1)) for hang protection. Every wait
ends on cancellation, so a time-limit overrun reports "Time limit was
exceeded" instead of hanging.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: bound lifecycle timeouts with a time limit, not a 5 s wall clock

keepalivePingIsBoundedWhenNoPongArrives and
handshakeTimeoutOptionBoundsTheWholeHandshake failed on PR #22's macOS job
(Xcode 27) with `ContinuousClock.now - start < .seconds(5)` at ~5.7 s. A
trivial test in the same window took 5.3 s, so the runner stalled; the
timeouts under test were 50 ms and 100 ms.

- keepalive ping: the socket never pongs, so the thrown URLError already
  proves the ping deadline ended the wait; a time limit catches a hang.
- handshake option: the fallback budget is raised to an hour through
  _test_setConnectTimeoutSeconds, so a channel that ignored the 100 ms option
  would trip the time limit instead of finishing at the 30 s default.
  Checked with a scratch test that drops the option: it fails with "Time
  limit was exceeded" and does not hang.

Both tests take .timeLimit(.minutes(1)) and pass while every cooperative
thread is blocked for several seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: show the device-auth write leaves the actor free by ordering, not elapsed time

issuedTokenPersistenceNeverBlocksTheChannelActor asserted the actor answered
within 3 s while a token write waited on another connection's SQLite lock.
A saturated test pool can stall the run for longer than that, and a time
limit alone would miss the regression it guards (a write on the actor holds
it for SQLite's 30 s busy timeout, less than the one-minute limit).

The test now relies on ordering. It releases the lock only after the actor
answers, so the token can reach disk only if the actor answered while the
write was pending. A write on the actor fails with SQLITE_BUSY first, the
token never lands, and the final wait trips the new suite time limit.

The 200 ms sleep that let hello-ok reach the write is replaced by a DEBUG
hook, _test_setDeviceTokenPersistenceStartedHandler, that fires on the actor
just before the persistence hop. Once the hop starts, shutdown cannot stop it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: bound the FIFO, AsyncTimeout and link-preview tests with a time limit

The three tests asserted wall-clock bounds (2 s, 5 s, 1 s) that a saturated
test pool can overrun. Each now has a one-minute time limit instead, and
each makes sure a regression ends on the limit's cancellation instead of
hanging the run:

- file fetch refuses a FIFO: a blocking open(2) never returns. The
  cancellation handler opens the FIFO's write end once to release it.
- operationThatIgnoresCancellationStillTimesOut: the parked operation is a
  gate the test opens on cancellation (and afterwards), not a
  never-resumed continuation that a loser-joining race would wait on
  forever.
- total deadline can fire before the session starts: URLSession's own
  timeouts move past the limit, so only the fetcher's zero-second
  deadline can end the fetch. The fetch is awaited through a no-deadline
  AsyncTimeout race, because a deadline lost before start would leave it
  unresumed even on cancellation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: drop wall-clock deadlines from the gateway and session-action waits

gatewayCoreWaitUntil gave up after 10 s (15 s at two call sites) and the
ChatViewModelSessionActionTests helpers after 15 s. A pool stall adds to
those waits, so they could time out even though the condition was about to
hold.

Both now wait until the condition holds or the test is cancelled. Every
suite that uses them has a one-minute time limit. gatewayCoreWaitUntil
records its GatewayCoreWaitTimeout as an issue before throwing, because
Swift Testing drops errors thrown after a time-limit cancellation and the
label says which wait hung. waitForForkStart awaits the gate's stream
directly instead of racing it against a sleep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: drop the wall-clock deadline from waitUntil

waitUntil gave up after 15 s (7, 10 or 30 s at five call sites). The macOS
CI job runs about 3,150 tests in parallel and the cooperative pool stalls
for 5 to 8 s at a time, so a wait could time out even though the condition
was about to hold.

waitUntil now polls until the condition holds or the test is cancelled, and
the timeoutSeconds and now: parameters are gone (nothing injected a clock,
and with no deadline there is nothing to measure). Every suite whose tests
reach waitUntil, directly or through a file-local helper, has a one-minute
time limit; ChatViewModelSessionActionTests already had one. On cancellation
the helper records AsyncWaitTimeoutError as an issue before throwing it,
because Swift Testing drops errors thrown after a time-limit cancellation
and the label says which wait hung.

No call site relied on the timeout: nothing expects AsyncWaitTimeoutError,
no wait runs in a child task that the test cancels, and the one try? wait is
cleanup in a catch block that rethrows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(e2e): wait without a wall-clock deadline in the channel and reconnect tests

ChannelAdaptersE2ETests.waitFor gave up after 15 s and only recorded an
unlabelled expectation failure. reconnectFailureSchedulesAnotherAttempt
polled against a 10 s deadline. A pool stall can outlast either one.

waitFor now takes a label, polls until the condition holds or the test is
cancelled, and records a WaitTimeout issue before throwing it. The reconnect
loop polls until cancelled. The suite and the reconnect test each have a
one-minute time limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: drop the 1 s deadline from the chat UI view-model waits

OpenClawChatUITests had its own waitUntil that gave up after 1 s and
returned false. On the macOS CI job for this PR, a pool stall of about
10 s ran past it: both view-model tests failed after about 12 s, and the
bootstrap expectations that followed failed with them.

The tests now call the shared deadline-free waitUntil with a label for
each wait, and the suite has a one-minute time limit. The private helper
is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: drop the wall-clock waits left in the provider and realtime relay tests

- ProviderStreamingCancellationTests polls with the shared deadline-free
  waitUntil under a one-minute suite time limit. Both request timeouts are
  now an hour, so a cancellation that never reaches the request fails as a
  hang instead of passing once the 60 s request timeout fires.
- ModelRouterStreamingFallbackTests waits for the stream termination (and
  the two stream/generate starts) with waitUntil, under a suite time limit.
- RealtimeTalkRelaySession._test_waitForStartupCancelled() drops its
  timeoutSeconds parameter. A closed session answers before any timer is
  armed; a zero timeout makes one that is not closed fail at once rather
  than after a 1 s wall-clock wait.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(linux): wait without a wall-clock deadline in the runtime tests

- A shared TestAsyncHelpers.swift replaces AgentLoopHardeningTests'
  waitUntil (500 x 10 ms, recorded an unlabelled issue and returned). It
  polls until the test is cancelled, then records and throws
  AsyncWaitTimeoutError(label:).
- GatewayServerTestHarness.collect (5 s) and Recorder.waitFor (5 s) take a
  label and wait until the time limit. The one deliberate negative wait now
  uses frames(_:arrivingWithinMs:), which returns what arrived; the old
  helper returned [] whenever its timer won, so that assertion could never
  fail.
- ChannelAdaptersLinuxSmokeTests.poll (15 s), the SIWC loopback-page wait
  (10 s), the automation run waits (10 s) and the MCP list_changed waits
  (3 s + 2 s) use waitUntil.
- addingAJobWakesTheSleepingLoop pushes the scheduler's idle cap to an hour
  through the new DEBUG hook CronScheduler._test_setMaximumSleepSeconds, so
  a missed wake hangs instead of racing the 60 s cap against the time limit.
- The MCP list_changed test uses an hour request timeout and drops its
  "< 1.5 s" elapsed assertion: a stall behind another request now hangs.
- .timeLimit(.minutes(1)) is on the nine suites that reach these waits and
  had no limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: drop the bounded polls, elapsed bounds and positive-path timeouts

#26 left three kinds of wall-clock bounds in the test waits. A stalled
cooperative pool (5-8 s on the macOS CI job) could fail any of them.

- Iteration-count sleep polls now use waitUntil, under a one-minute time
  limit. The Task.yield loops before negative checks, the sampler
  repetitions and the media teardown grace are kept.
- Elapsed "< N s" asserts are replaced by what they stood for: error
  payloads that name the short deadline, kill checks, and a stalled
  server or run that outlives the time limit. Lower bounds and the
  synchronous cron search are kept. The throttle drop and delay checks
  use an hour-long window, so a stall between the two requests cannot
  let the window pass.
- Positive-path product timeouts move past the time limit where the
  wait ends on cancellation (SIWC signIn, runtime.run). Waits that park
  on a continuation cancellation never resumes (runtime.wait, agent.wait,
  the brokers, A2ATaskStore, imsg) go through the new
  awaitCancellable(_:) with no timeout.
- The stale-timer test retries until its first run beats its 200 ms
  timer. The run-id collision test holds its run on a gate instead of a
  3 s sleep. The coalescing reporter test uses a frozen clock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: wait for agent runs without a positive-path timeout

#29 left runtime.wait and agent.wait calls with 2-5 s timeouts in the
agent gateway, session branch, loop, event stream and wire shape tests,
and in the OpenClawKitTests stack, gateway server and registry tests. A
stalled cooperative pool (5-8 s on the macOS CI job) can let that timer
beat the run and turn a passing wait into a "timeout".

Both waits park on a continuation that cancellation never resumes, so
they now go through awaitCancellable(_:) with no timeout, under a
one-minute time limit. OpenClawKitTests gets its own copy of the helper.
The client-side request timeout on the SDK gateway test's agent.wait
moves past the limit too. The intended short timeouts (timeoutMs 1 and
10) are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): 2026.3.2

Date the release 2026-10-01 and move the ConfigBoxed and config-decode stack
entries under it. Add a Tests section for the test-wait hardening (#22, #25,
#26, #29, #30), the per-test synthesizer (#23), the reply clock (#27) and the
stack-depth tests (#28), with the release test counts and CI gates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant