Repository navigation
fleet spawn reports failure (unconfirmed dispatch) for agents that did launch - #1942
agent-relay-code[bot] wants to merge 7 commits into
Conversation
`fleet spawn-status` was registered without updating the surfaces that enumerate commands, so the bootstrap leaf inventory and the Fleet CLI inventory test failed. - Add the leaf to `expectedLeafCommands` and to the current-main Fleet inventory counts, with an assertion that it is a public leaf: the spawn timeout error names the command, so a hidden or missing one makes that guidance unactionable. - Refresh the trusted Fleet CLI inventory snapshot and the matrix pin for the new leaf and the 360000ms `--confirm-timeout` default, and declare its option coverage. The Daytona board defers the leaf rather than mapping an operation: a read-only poll of a recorded invocation id proves nothing about placement that the spawn it reads has not already proven. Six `node` records with pre-existing snapshot drift (`--state-dir`, `--broker-url`, `--api-key`) are restored verbatim to keep that drift out of this change. - Cover the evidence the fix exists for: launch proof and the answering node id survive the sanitized read, and a dispatch that has not reported acquires no launch claim. - Document the command in the CLI README and the feature manifest, and split the changelog entry (new command is Added, hence Minor). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 99c51b3. Configure here.
…hangelog sections) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cation node ids spawn-status passed the workspace key alongside its minted launcher token, which createAgentRelay rejects, and sent that token to the default gateway rather than the one that minted it. Read with the token and the resolved workspace origin only, as the sandbox spawn path does. normalizeActionInvocation dropped handler_node_id and dispatched_node_id, so spawn-status could never report the node that answered a dispatch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…are-garden-f6614fab # Conflicts: # packages/sdk/src/__tests__/relaycast-translate.test.ts

Fix late fleet spawn confirmation and provide dispatch status reads
fleet spawnpreviously exhausted its two-minute confirmation budget while agents were still launching, then suggested redispatching without confirmation. Launches taking 1–5 minutes now have a six-minute default budget in the CLI and SDK. Explicit timeout overrides and the existing verified-spawn minimum remain supported.Add
agent-relay fleet spawn-status <invocation-id>to read the original spawn invocation without dispatching another worker. Timeout errors retain their structured invocation ID and direct callers to this command. It uses existing lifecycle receipts to report confirmed readiness, terminal failure, or an outcome still awaiting confirmation. A successful unverified launch retainsspawned: true, ready: falsewith stateaccepted. Workspace-only callers mint and clean up a temporary reader using the established launcher pattern. Output excludes task payloads and raw handler errors.Silence cannot establish that a process never started. A node that never returns a result remains explicitly unconfirmed; a terminal node failure is reported as failed. No roster-only inference of readiness is introduced. These tests use deterministic fixtures, not the affected physical fleet nodes.
The new leaf is carried onto the surfaces that enumerate commands: the bootstrap leaf-command inventory, the trusted Fleet CLI inventory snapshot (
tests/relayflows/cleanroom/fleet-cli-inventory.json) and its matrix pin, thefleetREADME, and the feature manifest. The Daytona board defersfleet spawn-statusrather than mapping it to an operation: a read-only poll of a recorded invocation id proves nothing about placement that the spawn it reads has not already proven.Validation:
sh .relayflow/check.sh— all checks passed (npm ci, codegen check, build, typecheck, lint, format, PR-proof and subscription guards,vitest run: 200 test files, 3754 tests).output.spawned/ready, dispatch and handler node ids) surviving the sanitized read; safe receipt output; temporary reader cleanup;fleet spawn-statuspresent as a public leaf in both command inventories.Two check failures were environmental, not code:
node-claim.test.tsneedslsof/psandbroker-process-identity.test.tscompiles a C fixture withcc. GitHub's ubuntu runners ship these (node-compat.ymlinstallslsof procps;test-install.ymlinstallsbuild-essential), so the install step was added to the uncommitted.relayflow/check.shrather than changing any test.Pre-existing, left alone: the committed Fleet CLI inventory snapshot is stale for six
nodecommands (--state-dir,--broker-url,--api-keywere never snapshotted afteraddBrokerOptionswas applied to them), so the qualification job's inventory comparison fails on main independently of this change. The snapshot regenerated here restores those six records verbatim to keep that drift out of this change.No workflow files changed. The existing unrelated active Trail trajectory prevented starting a new one; it was left intact.
Checks
Relayflow ran this repository's checks (.relayflow/check.sh) and they passed.
What ran (.relayflow/check.sh)
Fixes #1935
Note
Medium Risk
Changes fleet spawn confirmation timing and adds a new operational path for pending dispatches; mis-timed timeouts or misread poll states could still lead to duplicate spawns if operators ignore guidance, but behavior is read-only and backward-compatible for explicit timeouts.
Overview
Fixes late fleet spawn confirmation by extending the default launch-confirmation window from two to six minutes in the CLI (
fleet spawn --confirm-timeout) and SDK placement polling, so workers that register minutes after dispatch are confirmed instead of surfacingspawn_unconfirmed.Adds
agent-relay fleet spawn-status <invocation-id>, a read-only poll of the original spawn invocation (no second dispatch). JSON output uses sanitized lifecycle receipts withplacement.state(ready,failed,accepted,unconfirmed_may_be_running) and preserves dispatch/handler node evidence; workspace-only callers get a short-lived reader identity. Timeout errors now point at this command rather than suggesting a blind retry.The SDK
RelayActionInvocationtype andnormalizeActionInvocationnow surfacehandlerNodeId/dispatchedNodeIdwhen the server reports them. Command inventories, README, changelog, and feature manifest are updated accordingly.Reviewed by Cursor Bugbot for commit 3a6f0c0. Bugbot is set up for automated code reviews on this repo. Configure here.