Repository navigation
test: wait for agent runs without a positive-path timeout - #30
Closed
MarcoDotIO wants to merge 16 commits into
Closed
MarcoDotIO wants to merge 16 commits into
MarcoDotIO wants to merge 16 commits into
Conversation
anotherAgentsRunBeforeTheAckNeverHijacksTheIntent and preAckEventsOfTheAckedRunAreReplayed failed on main (CI run 36644330107, xcode-27 runner) with "Caught error: CancellationError()" after ~8 s; the same tree passed on PR #20. waitForSend polled the requester every 5 ms and threw once a 5 s ContinuousClock deadline passed. With ~3,150 Swift Testing tests saturating the cooperative pool, the host.send task (and the poller itself) did not run for more than 5 s, so the deadline expired although nothing in GatewayOpenClawIntentHost.send is slow. Blocking every cooperative thread for 6 s reproduces the failure locally. - OpenClawAppIntentsRunMatchingTests: HeldChatSendRequester yields each chat.send's params to an AsyncStream the moment the request arrives, and waitForSend awaits that stream. Assertions are unchanged. - GatewayNetworkConnectionTransportTests: the loopback NWListener start waits on stateUpdateHandler (ready/failed/cancelled) instead of polling listener.state against the same 5 s deadline, which also had no final re-check after a late wake-up. - WatchNodeClientTests: eventually() drops its 5 s deadline. Its conditions read production client state that has no change hook, so it still polls, but slowness no longer fails it. Each suite gains .timeLimit(.minutes(1)) for hang protection. Every wait ends on cancellation, so a time-limit overrun reports "Time limit was exceeded" instead of hanging. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
keepalivePingIsBoundedWhenNoPongArrives and handshakeTimeoutOptionBoundsTheWholeHandshake failed on PR #22's macOS job (Xcode 27) with `ContinuousClock.now - start < .seconds(5)` at ~5.7 s. A trivial test in the same window took 5.3 s, so the runner stalled; the timeouts under test were 50 ms and 100 ms. - keepalive ping: the socket never pongs, so the thrown URLError already proves the ping deadline ended the wait; a time limit catches a hang. - handshake option: the fallback budget is raised to an hour through _test_setConnectTimeoutSeconds, so a channel that ignored the 100 ms option would trip the time limit instead of finishing at the 30 s default. Checked with a scratch test that drops the option: it fails with "Time limit was exceeded" and does not hang. Both tests take .timeLimit(.minutes(1)) and pass while every cooperative thread is blocked for several seconds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ot elapsed time issuedTokenPersistenceNeverBlocksTheChannelActor asserted the actor answered within 3 s while a token write waited on another connection's SQLite lock. A saturated test pool can stall the run for longer than that, and a time limit alone would miss the regression it guards (a write on the actor holds it for SQLite's 30 s busy timeout, less than the one-minute limit). The test now relies on ordering. It releases the lock only after the actor answers, so the token can reach disk only if the actor answered while the write was pending. A write on the actor fails with SQLITE_BUSY first, the token never lands, and the final wait trips the new suite time limit. The 200 ms sleep that let hello-ok reach the write is replaced by a DEBUG hook, _test_setDeviceTokenPersistenceStartedHandler, that fires on the actor just before the persistence hop. Once the hop starts, shutdown cannot stop it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… limit The three tests asserted wall-clock bounds (2 s, 5 s, 1 s) that a saturated test pool can overrun. Each now has a one-minute time limit instead, and each makes sure a regression ends on the limit's cancellation instead of hanging the run: - file fetch refuses a FIFO: a blocking open(2) never returns. The cancellation handler opens the FIFO's write end once to release it. - operationThatIgnoresCancellationStillTimesOut: the parked operation is a gate the test opens on cancellation (and afterwards), not a never-resumed continuation that a loser-joining race would wait on forever. - total deadline can fire before the session starts: URLSession's own timeouts move past the limit, so only the fetcher's zero-second deadline can end the fetch. The fetch is awaited through a no-deadline AsyncTimeout race, because a deadline lost before start would leave it unresumed even on cancellation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aits gatewayCoreWaitUntil gave up after 10 s (15 s at two call sites) and the ChatViewModelSessionActionTests helpers after 15 s. A pool stall adds to those waits, so they could time out even though the condition was about to hold. Both now wait until the condition holds or the test is cancelled. Every suite that uses them has a one-minute time limit. gatewayCoreWaitUntil records its GatewayCoreWaitTimeout as an issue before throwing, because Swift Testing drops errors thrown after a time-limit cancellation and the label says which wait hung. waitForForkStart awaits the gate's stream directly instead of racing it against a sleep. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
waitUntil gave up after 15 s (7, 10 or 30 s at five call sites). The macOS CI job runs about 3,150 tests in parallel and the cooperative pool stalls for 5 to 8 s at a time, so a wait could time out even though the condition was about to hold. waitUntil now polls until the condition holds or the test is cancelled, and the timeoutSeconds and now: parameters are gone (nothing injected a clock, and with no deadline there is nothing to measure). Every suite whose tests reach waitUntil, directly or through a file-local helper, has a one-minute time limit; ChatViewModelSessionActionTests already had one. On cancellation the helper records AsyncWaitTimeoutError as an issue before throwing it, because Swift Testing drops errors thrown after a time-limit cancellation and the label says which wait hung. No call site relied on the timeout: nothing expects AsyncWaitTimeoutError, no wait runs in a child task that the test cancels, and the one try? wait is cleanup in a catch block that rethrows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nnect tests ChannelAdaptersE2ETests.waitFor gave up after 15 s and only recorded an unlabelled expectation failure. reconnectFailureSchedulesAnotherAttempt polled against a 10 s deadline. A pool stall can outlast either one. waitFor now takes a label, polls until the condition holds or the test is cancelled, and records a WaitTimeout issue before throwing it. The reconnect loop polls until cancelled. The suite and the reconnect test each have a one-minute time limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OpenClawChatUITests had its own waitUntil that gave up after 1 s and returned false. On the macOS CI job for this PR, a pool stall of about 10 s ran past it: both view-model tests failed after about 12 s, and the bootstrap expectations that followed failed with them. The tests now call the shared deadline-free waitUntil with a label for each wait, and the suite has a one-minute time limit. The private helper is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ay tests - ProviderStreamingCancellationTests polls with the shared deadline-free waitUntil under a one-minute suite time limit. Both request timeouts are now an hour, so a cancellation that never reaches the request fails as a hang instead of passing once the 60 s request timeout fires. - ModelRouterStreamingFallbackTests waits for the stream termination (and the two stream/generate starts) with waitUntil, under a suite time limit. - RealtimeTalkRelaySession._test_waitForStartupCancelled() drops its timeoutSeconds parameter. A closed session answers before any timer is armed; a zero timeout makes one that is not closed fail at once rather than after a 1 s wall-clock wait. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A shared TestAsyncHelpers.swift replaces AgentLoopHardeningTests' waitUntil (500 x 10 ms, recorded an unlabelled issue and returned). It polls until the test is cancelled, then records and throws AsyncWaitTimeoutError(label:). - GatewayServerTestHarness.collect (5 s) and Recorder.waitFor (5 s) take a label and wait until the time limit. The one deliberate negative wait now uses frames(_:arrivingWithinMs:), which returns what arrived; the old helper returned [] whenever its timer won, so that assertion could never fail. - ChannelAdaptersLinuxSmokeTests.poll (15 s), the SIWC loopback-page wait (10 s), the automation run waits (10 s) and the MCP list_changed waits (3 s + 2 s) use waitUntil. - addingAJobWakesTheSleepingLoop pushes the scheduler's idle cap to an hour through the new DEBUG hook CronScheduler._test_setMaximumSleepSeconds, so a missed wake hangs instead of racing the 60 s cap against the time limit. - The MCP list_changed test uses an hour request timeout and drops its "< 1.5 s" elapsed assertion: a stall behind another request now hangs. - .timeLimit(.minutes(1)) is on the nine suites that reach these waits and had no limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#26 left three kinds of wall-clock bounds in the test waits. A stalled cooperative pool (5-8 s on the macOS CI job) could fail any of them. - Iteration-count sleep polls now use waitUntil, under a one-minute time limit. The Task.yield loops before negative checks, the sampler repetitions and the media teardown grace are kept. - Elapsed "< N s" asserts are replaced by what they stood for: error payloads that name the short deadline, kill checks, and a stalled server or run that outlives the time limit. Lower bounds and the synchronous cron search are kept. The throttle drop and delay checks use an hour-long window, so a stall between the two requests cannot let the window pass. - Positive-path product timeouts move past the time limit where the wait ends on cancellation (SIWC signIn, runtime.run). Waits that park on a continuation cancellation never resumes (runtime.wait, agent.wait, the brokers, A2ATaskStore, imsg) go through the new awaitCancellable(_:) with no timeout. - The stale-timer test retries until its first run beats its 200 ms timer. The run-id collision test holds its run on a gate instead of a 3 s sleep. The coalescing reporter test uses a frozen clock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#29 left runtime.wait and agent.wait calls with 2-5 s timeouts in the agent gateway, session branch, loop, event stream and wire shape tests, and in the OpenClawKitTests stack, gateway server and registry tests. A stalled cooperative pool (5-8 s on the macOS CI job) can let that timer beat the run and turn a passing wait into a "timeout". Both waits park on a continuation that cancellation never resumes, so they now go through awaitCancellable(_:) with no timeout, under a one-minute time limit. OpenClawKitTests gets its own copy of the helper. The client-side request timeout on the SDK gateway test's agent.wait moves past the limit too. The intended short timeouts (timeoutMs 1 and 10) are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…meout Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d-elapsed-asserts
2 of 3 tasks
MarcoDotIO
added a commit
that referenced
this pull request
Oct 1, 2026
* test: wait for events instead of a 5 s wall-clock deadline anotherAgentsRunBeforeTheAckNeverHijacksTheIntent and preAckEventsOfTheAckedRunAreReplayed failed on main (CI run 36644330107, xcode-27 runner) with "Caught error: CancellationError()" after ~8 s; the same tree passed on PR #20. waitForSend polled the requester every 5 ms and threw once a 5 s ContinuousClock deadline passed. With ~3,150 Swift Testing tests saturating the cooperative pool, the host.send task (and the poller itself) did not run for more than 5 s, so the deadline expired although nothing in GatewayOpenClawIntentHost.send is slow. Blocking every cooperative thread for 6 s reproduces the failure locally. - OpenClawAppIntentsRunMatchingTests: HeldChatSendRequester yields each chat.send's params to an AsyncStream the moment the request arrives, and waitForSend awaits that stream. Assertions are unchanged. - GatewayNetworkConnectionTransportTests: the loopback NWListener start waits on stateUpdateHandler (ready/failed/cancelled) instead of polling listener.state against the same 5 s deadline, which also had no final re-check after a late wake-up. - WatchNodeClientTests: eventually() drops its 5 s deadline. Its conditions read production client state that has no change hook, so it still polls, but slowness no longer fails it. Each suite gains .timeLimit(.minutes(1)) for hang protection. Every wait ends on cancellation, so a time-limit overrun reports "Time limit was exceeded" instead of hanging. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: bound lifecycle timeouts with a time limit, not a 5 s wall clock keepalivePingIsBoundedWhenNoPongArrives and handshakeTimeoutOptionBoundsTheWholeHandshake failed on PR #22's macOS job (Xcode 27) with `ContinuousClock.now - start < .seconds(5)` at ~5.7 s. A trivial test in the same window took 5.3 s, so the runner stalled; the timeouts under test were 50 ms and 100 ms. - keepalive ping: the socket never pongs, so the thrown URLError already proves the ping deadline ended the wait; a time limit catches a hang. - handshake option: the fallback budget is raised to an hour through _test_setConnectTimeoutSeconds, so a channel that ignored the 100 ms option would trip the time limit instead of finishing at the 30 s default. Checked with a scratch test that drops the option: it fails with "Time limit was exceeded" and does not hang. Both tests take .timeLimit(.minutes(1)) and pass while every cooperative thread is blocked for several seconds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: show the device-auth write leaves the actor free by ordering, not elapsed time issuedTokenPersistenceNeverBlocksTheChannelActor asserted the actor answered within 3 s while a token write waited on another connection's SQLite lock. A saturated test pool can stall the run for longer than that, and a time limit alone would miss the regression it guards (a write on the actor holds it for SQLite's 30 s busy timeout, less than the one-minute limit). The test now relies on ordering. It releases the lock only after the actor answers, so the token can reach disk only if the actor answered while the write was pending. A write on the actor fails with SQLITE_BUSY first, the token never lands, and the final wait trips the new suite time limit. The 200 ms sleep that let hello-ok reach the write is replaced by a DEBUG hook, _test_setDeviceTokenPersistenceStartedHandler, that fires on the actor just before the persistence hop. Once the hop starts, shutdown cannot stop it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: bound the FIFO, AsyncTimeout and link-preview tests with a time limit The three tests asserted wall-clock bounds (2 s, 5 s, 1 s) that a saturated test pool can overrun. Each now has a one-minute time limit instead, and each makes sure a regression ends on the limit's cancellation instead of hanging the run: - file fetch refuses a FIFO: a blocking open(2) never returns. The cancellation handler opens the FIFO's write end once to release it. - operationThatIgnoresCancellationStillTimesOut: the parked operation is a gate the test opens on cancellation (and afterwards), not a never-resumed continuation that a loser-joining race would wait on forever. - total deadline can fire before the session starts: URLSession's own timeouts move past the limit, so only the fetcher's zero-second deadline can end the fetch. The fetch is awaited through a no-deadline AsyncTimeout race, because a deadline lost before start would leave it unresumed even on cancellation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: drop wall-clock deadlines from the gateway and session-action waits gatewayCoreWaitUntil gave up after 10 s (15 s at two call sites) and the ChatViewModelSessionActionTests helpers after 15 s. A pool stall adds to those waits, so they could time out even though the condition was about to hold. Both now wait until the condition holds or the test is cancelled. Every suite that uses them has a one-minute time limit. gatewayCoreWaitUntil records its GatewayCoreWaitTimeout as an issue before throwing, because Swift Testing drops errors thrown after a time-limit cancellation and the label says which wait hung. waitForForkStart awaits the gate's stream directly instead of racing it against a sleep. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: drop the wall-clock deadline from waitUntil waitUntil gave up after 15 s (7, 10 or 30 s at five call sites). The macOS CI job runs about 3,150 tests in parallel and the cooperative pool stalls for 5 to 8 s at a time, so a wait could time out even though the condition was about to hold. waitUntil now polls until the condition holds or the test is cancelled, and the timeoutSeconds and now: parameters are gone (nothing injected a clock, and with no deadline there is nothing to measure). Every suite whose tests reach waitUntil, directly or through a file-local helper, has a one-minute time limit; ChatViewModelSessionActionTests already had one. On cancellation the helper records AsyncWaitTimeoutError as an issue before throwing it, because Swift Testing drops errors thrown after a time-limit cancellation and the label says which wait hung. No call site relied on the timeout: nothing expects AsyncWaitTimeoutError, no wait runs in a child task that the test cancels, and the one try? wait is cleanup in a catch block that rethrows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(e2e): wait without a wall-clock deadline in the channel and reconnect tests ChannelAdaptersE2ETests.waitFor gave up after 15 s and only recorded an unlabelled expectation failure. reconnectFailureSchedulesAnotherAttempt polled against a 10 s deadline. A pool stall can outlast either one. waitFor now takes a label, polls until the condition holds or the test is cancelled, and records a WaitTimeout issue before throwing it. The reconnect loop polls until cancelled. The suite and the reconnect test each have a one-minute time limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: drop the 1 s deadline from the chat UI view-model waits OpenClawChatUITests had its own waitUntil that gave up after 1 s and returned false. On the macOS CI job for this PR, a pool stall of about 10 s ran past it: both view-model tests failed after about 12 s, and the bootstrap expectations that followed failed with them. The tests now call the shared deadline-free waitUntil with a label for each wait, and the suite has a one-minute time limit. The private helper is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: drop the wall-clock waits left in the provider and realtime relay tests - ProviderStreamingCancellationTests polls with the shared deadline-free waitUntil under a one-minute suite time limit. Both request timeouts are now an hour, so a cancellation that never reaches the request fails as a hang instead of passing once the 60 s request timeout fires. - ModelRouterStreamingFallbackTests waits for the stream termination (and the two stream/generate starts) with waitUntil, under a suite time limit. - RealtimeTalkRelaySession._test_waitForStartupCancelled() drops its timeoutSeconds parameter. A closed session answers before any timer is armed; a zero timeout makes one that is not closed fail at once rather than after a 1 s wall-clock wait. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(linux): wait without a wall-clock deadline in the runtime tests - A shared TestAsyncHelpers.swift replaces AgentLoopHardeningTests' waitUntil (500 x 10 ms, recorded an unlabelled issue and returned). It polls until the test is cancelled, then records and throws AsyncWaitTimeoutError(label:). - GatewayServerTestHarness.collect (5 s) and Recorder.waitFor (5 s) take a label and wait until the time limit. The one deliberate negative wait now uses frames(_:arrivingWithinMs:), which returns what arrived; the old helper returned [] whenever its timer won, so that assertion could never fail. - ChannelAdaptersLinuxSmokeTests.poll (15 s), the SIWC loopback-page wait (10 s), the automation run waits (10 s) and the MCP list_changed waits (3 s + 2 s) use waitUntil. - addingAJobWakesTheSleepingLoop pushes the scheduler's idle cap to an hour through the new DEBUG hook CronScheduler._test_setMaximumSleepSeconds, so a missed wake hangs instead of racing the 60 s cap against the time limit. - The MCP list_changed test uses an hour request timeout and drops its "< 1.5 s" elapsed assertion: a stall behind another request now hangs. - .timeLimit(.minutes(1)) is on the nine suites that reach these waits and had no limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: drop the bounded polls, elapsed bounds and positive-path timeouts #26 left three kinds of wall-clock bounds in the test waits. A stalled cooperative pool (5-8 s on the macOS CI job) could fail any of them. - Iteration-count sleep polls now use waitUntil, under a one-minute time limit. The Task.yield loops before negative checks, the sampler repetitions and the media teardown grace are kept. - Elapsed "< N s" asserts are replaced by what they stood for: error payloads that name the short deadline, kill checks, and a stalled server or run that outlives the time limit. Lower bounds and the synchronous cron search are kept. The throttle drop and delay checks use an hour-long window, so a stall between the two requests cannot let the window pass. - Positive-path product timeouts move past the time limit where the wait ends on cancellation (SIWC signIn, runtime.run). Waits that park on a continuation cancellation never resumes (runtime.wait, agent.wait, the brokers, A2ATaskStore, imsg) go through the new awaitCancellable(_:) with no timeout. - The stale-timer test retries until its first run beats its 200 ms timer. The run-id collision test holds its run on a gate instead of a 3 s sleep. The coalescing reporter test uses a frozen clock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: wait for agent runs without a positive-path timeout #29 left runtime.wait and agent.wait calls with 2-5 s timeouts in the agent gateway, session branch, loop, event stream and wire shape tests, and in the OpenClawKitTests stack, gateway server and registry tests. A stalled cooperative pool (5-8 s on the macOS CI job) can let that timer beat the run and turn a passing wait into a "timeout". Both waits park on a continuation that cancellation never resumes, so they now go through awaitCancellable(_:) with no timeout, under a one-minute time limit. OpenClawKitTests gets its own copy of the helper. The client-side request timeout on the SDK gateway test's agent.wait moves past the limit too. The intended short timeouts (timeoutMs 1 and 10) are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): 2026.3.2 Date the release 2026-10-01 and move the ConfigBoxed and config-decode stack entries under it. Add a Tests section for the test-wait hardening (#22, #25, #26, #29, #30), the per-test synthesizer (#23), the reply clock (#27) and the stack-depth tests (#28), with the release test counts and CI gates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
6 tasks done
Owner
Author
|
Landed on main in dca4976 via #31. #31's branch was main plus #30's head, and its squash-merge went in before this PR, so dca4976 carries this PR's changes and lists its commits. Merging this PR now would only add an empty commit, so I'm closing it as landed. Checked: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
#29 listed the positive-path
agent.wait/runtime.waitcalls with 2–5 s timeouts that it left out of scope. This PR converts them. The macOS CI job runs about 3,150 tests in parallel and can stall the cooperative pool for 5–8 s. A stall that lands while one of these runs is in flight lets the wait's own timer beat the run, and a passing test reportstimeout.Raising the timeout isn't enough.
EmbeddedAgentRuntime.wait(runID:timeoutMs:)and the gateway'shandleAgentWaitboth park on a rawCheckedContinuationthat only their timer or the finished run resumes, and cancellation never does. So each call now goes throughawaitCancellable(_ label:)with no timeout, under a one-minute.timeLimit. A regression shows up as a hang, and the time limit turns it into anAsyncWaitTimeoutErrorthat names the wait.OpenClawKitTests gets its own
awaitCancellable(plusrecordedWaitTimeout, which itswaitUntilnow uses too) inTests/OpenClawKitTests/TestAsyncHelpers.swift, the same as the Linux copy.AgentGatewayMethodsTests:51agent.wait5 s → runokawaitCancellable, no timeout; suite time limitAgentGatewayMethodsTests:88runtime.wait5 s →okAgentGatewayMethodsTests:218agent.wait2 s aftersessions.abort→errorSessionBranchGatewayMethodsTests:27, :127, :164runtime.wait5 s →okawaitCancellable, no timeout (the suite already had a time limit)AgentLoopToolCallingTests:292runtime.wait2 s afterabort→errorawaitCancellable, no timeout; per-test time limitGatewayEventStreamTests:195agent.wait2 s afterchat.abort→errorawaitCancellable, no timeout (suite already limited)GatewayWireShapeTests:255agent.wait2 s →okGatewayWireShapeTests:280agent.wait2 s on the 150 ms slow run →okEmbeddedAgentStackTests:35runtime.wait5 s → outputawaitCancellable, no timeout; suite time limitGatewayServerTests:197, :233agent.wait5 s / 1 s →okawaitCancellable, no timeout; per-test time limitGatewayServerTests:817GatewayClient.request("agent.wait")withtimeoutMs: 5_000→okGatewayClient.sendOncealso parks on a continuation cancellation never resumes.GatewayServerRegistryTests:464agent.wait1 s →okawaitCancellable, no timeout; per-test time limitGatewayServerRegistryTests:467agent.wait, no timeout, legacyrunIDkeyawaitCancellabletoo, so the new time limit can end itThe suite time limit goes on the small suites. The large ones (
AgentLoopToolCallingTestswith 15 tests, the.serializedGatewayServerTestsandGatewayServerRegistryTests) get it per test instead.Kept, on purpose. These are intended timeouts, or they answer from
completedRunswithout arming a timer:GatewayWireShapeTests:276,timeoutMs: 1(the timed-out wait under test)GatewayServerTests:207,timeoutMs: 10(late wait on a finished run)GatewayServerTests:225,timeoutMs: 1(the timed-out wait under test)AgentGatewayMethodsTests:166,question.waitAnswertimeoutMs: 10(expectspending)Not in this PR
These are still wall-clock windows in the same files, but they aren't
agent.wait/runtime.waittimeouts:sessionsAbortAndCompact,abortCancelsRunAndRecordsAbortedAssistant,abortEmitsAbortedChatEventAndFiltersSelectEventsandmutationsRefuseWhileARunIsActivehold the run in a 5 s scripted provider sleep and abort it after a 50–100 ms sleep. A stall of 5 s or more before the abort lets the run finish first.GatewayWireShapeTests:276,GatewayServerTests:225), the run only sleeps 150 ms / 80 ms. A stall before the wait arms its timer lets the run finish first, so the wait reportsokinstead oftimeout. AnAsyncGatethe test opens after the timed-out wait would remove that window.runtime.runtimeouts.AgentLoopToolCallingTestsstill callsruntime.run(timeoutMs:)with 10 s at :191 and 5 s at :341, :342 and :371.runtime.runends on cancellation, so these can move past the limit the same way test: drop the bounded polls, elapsed bounds and positive-path timeouts #29 did inSubagentRuntimeHardeningTests.GatewayClient.requestcalls insdkGatewayServerSupportsCatalogSkillsAndRuntimeHandlerskeep the 15 s default.Stacked on #29 (→ #26 → #25/#24). Merge those first; until then this diff also includes their commits. Only 07e589b is new. The merge of
mainafter it (5912d21) changes no files. Main has #25 squash-merged, while the stack carries #25's own commit, soTestAsyncHelpers.swiftneeded a resolution.Test plan
swift build --build-testsScripts/lint-swift.sh: 0 violations in 1,128 filesTargeted
swift test --filterover the 8 affected suites: 64 tests pass.Full
swift test: OpenClawKitTests 3,150 tests and OpenClawKitE2ETests 20 pass. In OpenClawLinuxRuntimeTests, 581 of 582 pass. The failure is the knownSpotlightMemoryTests.replyCoordinatorRoutesRepliesToOneInvocationAtATimeflake, which failed 1 of 6 isolated runs, and this PR doesn't touch it. Two more runs of the whole Linux bundle passed 582/582.Pool starvation. An uncommitted scratch hook inside both
awaitCancellablecopies spawnedactiveProcessorCount * 2Task.detached(priority: Task.currentPriority) { usleep(6_000_000) }blockers, gated onOPENCLAW_STARVE=1. That stalls the pool for about 12 s, at most 3 times per process. Each test ran in its ownswift test --filter Suite/testprocess.All 15 tests that reach a converted wait pass, after 12–36 s: 21 stalls in all, and each of the 17
awaitCancellablesites was stalled at least once.AgentGatewayMethodsTestssessionsSendRunsTheLoopAndHistoryReturnsAgentMessagesAgentGatewayMethodsTestssessionsCreateStartsARunAndPatchCancelsApprovalsOnPermissionChangeAgentGatewayMethodsTestssessionsAbortAndCompactSessionBranchGatewayMethodsTestsrewindMovesTheLeafAndBranchesListAndSwitchSessionBranchGatewayMethodsTestsforkCopiesTheActivePathBeforeTheEntrySessionBranchGatewayMethodsTestssearchMatchesWordsAndPhrasesPerAgentSessionBranchGatewayMethodsTestsmutationsRefuseWhileARunIsActiveAgentLoopToolCallingTestsabortCancelsRunAndRecordsAbortedAssistantGatewayEventStreamTestsabortEmitsAbortedChatEventAndFiltersSelectEventsGatewayWireShapeTestsagentAcceptsUpstreamParamsAndDedupesRetriesGatewayWireShapeTestsagentWaitTimeoutKeepsTrackingAndLegacyKeyWorksEmbeddedAgentStackTestsstackPersistsTranscriptsAndServesRuntimeRPCsGatewayServerTestsgatewayServerSupportsAgentRunWaitTimeoutAndCleanupGatewayServerTestssdkGatewayServerSupportsCatalogSkillsAndRuntimeHandlersGatewayServerRegistryTestsbuiltinHandlersAcceptUpstreamWireShapesControl. The hook sits inside
awaitCancellable, so the old code never reaches it. To compare the two, a second scratch hook stalled the pool 10 ms after the wait started, while the run was still in flight. It was applied to the two sites whose run lasts long enough to still be running at that point (the 150 ms and 80 ms slow runs).GatewayWireShapeTests:280 (150 ms run, oldtimeoutMs: 2_000)done["status"] → "timeout"GatewayServerTests:233 (80 ms run, oldtimeoutMs: 1_000)eventuallyCompleted.status → "timeout"This is the CI failure mode: the wait's timer fires during the stall and gets to the server before the run's completion does. The other sites' scripted runs finish within milliseconds of the wait starting. A CI stall that lands in that window has the same effect, but a scratch hook can't hit it reliably.
🤖 Generated with Claude Code