🔁 fix: Retry Code API Rate Limits That Reject Before a Workspace Call Starts - #16453
Conversation
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
LibreChat/packages/api/src/code/workspace.ts
Lines 940 to 942 in 3247167
When a proxy drops or mangles Retry-After, this passes only the missing header to workspaceAdmissionRetryDelay, so every typed 429 is retried after the one-second default instead of its retry_after_seconds value. The Code API body field is explicitly the proxy-safe fallback already supported by getCodeApiRetryAfterMs in packages/api/src/utils/code.ts; ignoring it here can generate repeated premature requests and exhaust the retry horizon before the limiter permits another call. Parse the buffered body’s delay when the header is unusable.
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Follow-up ready at |
|
Follow-up head |
|
Latest pushed head |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: de32233b99
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Review handoff for exact pushed head |
Summary
When several agents share one attached (BYOM) code environment, a burst of workspace calls trips the Code API's per-user rate limiter, which returns
429with a typed body ({"error":"rate_limited","retry_after_seconds":N}) and aRetry-Afterheader. LibreChat surfaced every one of those as a failed tool call ("Workspace tool request rejected (upstreamStatus: 429)"), so the model saw a hard failure and often gave up or improvised. On one deployment 18 of 72 workspace calls (25%) failed this way during a parallel-agent burst.The limiter runs as middleware ahead of the Code API's workspace router, so a typed 429 means the operation was never assigned or started, the same guarantee a
503 WORKSPACE_QUEUE_TIMEOUTalready carries. This change treats both as admission rejections, but bounds their retries separately: 503 uses the workspace queue horizon, while typed 429s use the existingendpoints.agents.codeApiMaxRetryWaitMsbudget (20 seconds by default;0disables rate-limit retries), even when capacity retries are disabled or exhausted. Both remain subject to the caller HTTP deadline and its reserved execution time. The client prefersRetry-After, falls back to the typed body’sretry_after_secondswhen a proxy drops or mangles that header, and caps each wait at 30 seconds. When a budget runs out, the error tells the model the operation was not started and to slow down. Anything else under 429 or 503, including an untyped or truncated body, keeps an unknown outcome and is never retried.Related to LibreChat-AI/code-interpreter#267 (BYOM defaults and parallel-agent limits).
How it works
Queue waits never consume the rate-limit wait budget, and an accepted rate-limit wait earns one retry even if its timer fires late. An unaffordable rate-limit hint fails immediately; later 429s cannot borrow time from a spent budget. The three request sources (Bash, file operations, and repository instructions) use the same request config. Each retry obtains fresh credentials; cached instructions still require authorization. A protected preview/edit retains the capacity retry horizon across both requests, but a successful preview always permits one hash-guarded edit attempt even when that horizon is spent. Untyped 429s and ambiguous failures are never retried.
Type of change
Testing
Tested environments/configuration:
code-interpretermain (337ddd60a8ea):service/src/middleware/limits.tsemits the typed 429 body andservice/src/workspace-tools/index.tsinstalls the limiter before the router. A live BYOM request was not rerun for this head.Automated tests:
packages/api: 523 passed acrosssrc/code/workspace.spec.ts,src/code/instructions.spec.ts,src/code/command.spec.ts,src/agents/handlers.spec.ts,src/agents/__tests__/initialize.test.ts, andsrc/utils/code.spec.ts. Coverage includes isolated 503/429 horizons, timer jitter, cancellation, HTTP execution reserve, malformed delay hints, input limits, authorization refresh, and protected edit preview-to-mutation transitions.api: 308 passed acrossserver/services/Files/Code/process.spec.jsandserver/services/__tests__/ToolService.spec.js, including both definitions-only and eager request-config propagation.packages/api: no-emit TypeScript check and production package build.npm run static-checks -- --against origin/devpassed all affected PR-wide gates, including the circular-dependency scan. Local Lighthouse was not run because Chrome is unavailable; the PR has a Lighthouse CI lane.