Repository navigation
CLI: one-shot 'relayfile read' does not retry 429 workspace_busy (Retry-After is parsed but unused) - #527
CLI: one-shot 'relayfile read' does not retry 429 workspace_busy (Retry-After is parsed but unused)#527agent-relay-code[bot] wants to merge 4 commits into
Conversation
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Devin Review found 3 potential issues.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| } | ||
| var apiErr *apiError | ||
| _ = errors.As(lastErr, &apiErr) | ||
| return pollResult{err: lastErr, httpStatus: apiErr.StatusCode, retryAfter: apiErr.RetryAfter} |
There was a problem hiding this comment.
🔴 Busy workspace retries too early
For a 429 workspace_busy response, politePoll can shorten the advertised Retry-After with negative jitter. The next request can hit the still-busy workspace and exhaust the retry budget.
Learn more
The one-shot retry passes the parsed server delay to politePoll. That poller clamps the delay and then applies positive or negative jitter. Negative jitter sends a request before a busy workspace's advertised recovery time, even though the retry is intended to respect the server's hint.
Example: A workspace answers 429 with Retry-After: 7. The poller's 20% jitter can schedule the next request after 5.6 seconds instead of at least seven seconds. Repeated early retries can use all three attempts before the workspace is ready.
Recommended fix: Treat Retry-After as a minimum wait when scheduling one-shot retries. Apply jitter only as extra time, or take the maximum of the server delay and the jittered client backoff, while retaining the intended upper bound.
Was this helpful? React with 👍 or 👎 to provide feedback.
| } | ||
| var apiErr *apiError | ||
| _ = errors.As(lastErr, &apiErr) | ||
| return pollResult{err: lastErr, httpStatus: apiErr.StatusCode, retryAfter: apiErr.RetryAfter} |
There was a problem hiding this comment.
🔴 Service retries ignore recovery time
For a 503 with Retry-After, politePoll ignores the server delay and retries after its one-second fallback. All three attempts can finish before the service's advertised recovery time.
Learn more
The one-shot policy retries both 429 and 503 errors. It forwards the parsed delay to politePoll, but that loop applies retryAfter only when httpStatus == 429. A 503 always takes exponential backoff regardless of its header.
Example: A server responds 503 with Retry-After: 30. The CLI retries at roughly one and then three seconds, returns the final 503, and never waits for the advertised 30-second recovery.
Recommended fix: Extend the wait calculation for one-shot requests to honor valid Retry-After on 503 as well as 429, preserving the bounded delay and retry count.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if errors.As(err, &apiErr) && apiErr.StatusCode == http.StatusTooManyRequests && apiErr.Code == "workspace_busy" { | ||
| return "workspace busy" | ||
| } | ||
| return "service unavailable" |
There was a problem hiding this comment.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit cda02fb. Configure here.
Relayfile Eval ReviewRun: Passed: 4 | Needs human: 0 | Reviewable: 0 | Missing output: 0 | Failed: 0 | Skipped: 0 Human Review CasesNo reviewable human-review cases captured Relayfile output. |

PR summary
What changed
429and503responses to the one-shottree,read,export, andstatusGET commands.Retry-After, exponential fallback, jitter, rate limiting, and context-aware waits, with three attempts and a 60-second maximum delay.--no-retryto each affected command for callers that require immediate failure.Tests
429 workspace_busyrecovery,503, exhaustion, non-retryable errors, delay clamping, cancellation, stdout/stderr separation, andread --no-retryrequest count.go test ./cmd/relayfile-cli -count=1scripts/check-contract-surface.shgit diff --checkChecks
Relayflow ran this repository's checks (.relayflow/check.sh) and they passed.
What ran (.relayflow/check.sh)
Fixes #526
Note
Low Risk
CLI-only resilience for read paths; default behavior adds waits on transient errors but preserves prior failure modes via --no-retry.
Overview
Adds automatic bounded retries for transient
429/503responses on one-shot workspace GET flows (tree,read,export,status, andpull’s provider discovery via sync status).A new
oneShotGETRetryPolicywraps workspace JSON/byte reads with up to three attempts, honoringRetry-After, exponential backoff with jitter (capped at 60s), and human-readable wait lines on stderr so stdout payloads stay clean.--no-retryon each affected command restores immediate failure for scripts that need it.Command handlers and the published CLI spec/docs are updated to document the flag and default retry behavior; regression tests cover
workspace_busy, delay clamping, cancellation, stderr separation, and--no-retry.Reviewed by Cursor Bugbot for commit 2b1bd9b. Bugbot is set up for automated code reviews on this repo. Configure here.
Agent Relay sessions
claudesession03e53226· contributor · rangh prcommands · last active 2026-10-05