With several agents running in parallel against one BYOM worker (4 negotiated lease slots, six registered repositories), 503 COMMAND_UNAVAILABLE is now the largest remaining source of failed workspace calls: 11 of 680 calls in one afternoon, after rate limits (#267) and queue timeouts were addressed. Clients cannot safely retry it, because the code is shared by failures with very different outcomes.
One code, three meanings
| Where |
Condition |
Did the operation start? |
packages/code/src/native-pool.ts:122-126 (allocate) |
Native workspace already executing: the root's executor entry is busy |
No |
native-pool.ts:128-136 |
Native executor capacity reached: every pool entry is busy |
No |
native-pool.ts:151-159 |
Allocation or idle-eviction failure (the comment says it precedes dispatch) |
No |
native-sandbox.ts:547-551 (execute) |
already has an active command or is closing |
No |
native-process.ts:556 / :460 |
Native executor setup failed before dispatch (the message we observe) |
No |
workspace-runtime.ts:121-126 |
Any non-400 4xx from the sandboxed runtime is mapped to COMMAND_UNAVAILABLE |
Unknown |
workspace-runtime.ts:131-140 |
Runtime transport failure; the session is quarantined, then COMMAND_UNAVAILABLE |
Unknown |
native-sandbox.ts:801, :820 |
Probe network isolation unavailable / cleanup failed |
Depends on phase |
Most occurrences are "not started", but because some are not, a client has to treat every COMMAND_UNAVAILABLE as an unknown outcome. LibreChat now retries 503 WORKSPACE_QUEUE_TIMEOUT and typed 429 rate_limited inside the caller's budget (LibreChat-AI/LibreChat#16453), and deliberately does not retry this code.
The busy case should not happen after admission
The Code API only admits one assignment per workspace isolation key, and the worker tracks active assignments by the same key (worker.ts activeWorkspaceAssignments). So by the time a workspace-tool assignment reaches NativeExecutorPool.allocate, its root should not be busy. That it is suggests two admission paths that do not share one lock over the same pool root. The likeliest candidate is the programmatic tool-calling path (native-programmatic.ts, "Native programmatic executor") running on the same root as a workspace-tool assignment, or an executor that is still settling (cleanup or output upload) after its assignment was released. The pool's capacity (1–8) and the negotiated lease slots are also separate limits; if the pool is smaller than the slot count, admitted work can find every entry busy.
Proposal
- Wait instead of refusing inside an admitted assignment. If the root's executor is busy or the pool is full, wait for it within the assignment's remaining execution budget (honoring cancellation), instead of throwing. Admission already bounded concurrency; a second refusal at execution time only converts contention into failures.
- Split the code. Keep
COMMAND_UNAVAILABLE for genuine unavailability, and return a distinct typed code for pre-dispatch refusals that guarantees nothing started, for example WORKSPACE_EXECUTOR_BUSY (503, with Retry-After). Clients can then retry it safely the same way they retry WORKSPACE_QUEUE_TIMEOUT.
- Stop remapping unknown outcomes into it.
workspace-runtime.ts should keep a distinct code for runtime 4xx and transport failures, whose outcome is uncertain, so "unavailable" never hides a possibly started command.
- Tie pool capacity to the negotiated slots (at least one executor per slot), or document the relationship and warn at startup when the pool is smaller.
- Find the double admission. Add the isolation key and admission path (workspace tool or programmatic) to the busy error's log context, so the second path contending for a root can be identified.
Acceptance
- With 4 slots and 7 concurrent agents across repositories, no workspace call fails
COMMAND_UNAVAILABLE because an executor is busy; contention shows up as queueing, not errors.
- A pre-dispatch refusal returns a typed not-started code that a client can retry, and any response whose outcome is uncertain never uses that code.
- A regression test runs a programmatic execution and a workspace command against one root concurrently and asserts they serialize without either being refused.
Related: #267 (item 5), LibreChat-AI/LibreChat#16453.
With several agents running in parallel against one BYOM worker (4 negotiated lease slots, six registered repositories),
503 COMMAND_UNAVAILABLEis now the largest remaining source of failed workspace calls: 11 of 680 calls in one afternoon, after rate limits (#267) and queue timeouts were addressed. Clients cannot safely retry it, because the code is shared by failures with very different outcomes.One code, three meanings
packages/code/src/native-pool.ts:122-126(allocate)Native workspace already executing: the root's executor entry isbusynative-pool.ts:128-136Native executor capacity reached: every pool entry is busynative-pool.ts:151-159native-sandbox.ts:547-551(execute)already has an active command or is closingnative-process.ts:556/:460Native executor setup failed before dispatch(the message we observe)workspace-runtime.ts:121-126COMMAND_UNAVAILABLEworkspace-runtime.ts:131-140COMMAND_UNAVAILABLEnative-sandbox.ts:801,:820Most occurrences are "not started", but because some are not, a client has to treat every
COMMAND_UNAVAILABLEas an unknown outcome. LibreChat now retries503 WORKSPACE_QUEUE_TIMEOUTand typed429 rate_limitedinside the caller's budget (LibreChat-AI/LibreChat#16453), and deliberately does not retry this code.The busy case should not happen after admission
The Code API only admits one assignment per workspace isolation key, and the worker tracks active assignments by the same key (
worker.tsactiveWorkspaceAssignments). So by the time a workspace-tool assignment reachesNativeExecutorPool.allocate, its root should not bebusy. That it is suggests two admission paths that do not share one lock over the same pool root. The likeliest candidate is the programmatic tool-calling path (native-programmatic.ts, "Native programmatic executor") running on the same root as a workspace-tool assignment, or an executor that is still settling (cleanup or output upload) after its assignment was released. The pool'scapacity(1–8) and the negotiated lease slots are also separate limits; if the pool is smaller than the slot count, admitted work can find every entry busy.Proposal
COMMAND_UNAVAILABLEfor genuine unavailability, and return a distinct typed code for pre-dispatch refusals that guarantees nothing started, for exampleWORKSPACE_EXECUTOR_BUSY(503, withRetry-After). Clients can then retry it safely the same way they retryWORKSPACE_QUEUE_TIMEOUT.workspace-runtime.tsshould keep a distinct code for runtime 4xx and transport failures, whose outcome is uncertain, so "unavailable" never hides a possibly started command.Acceptance
COMMAND_UNAVAILABLEbecause an executor is busy; contention shows up as queueing, not errors.Related: #267 (item 5), LibreChat-AI/LibreChat#16453.