Description
The MCP initialize handshake borrows the steady-state per-request timeout (request_timeout_seconds) as its start-up budget. A slow but healthy server (heavy module import at spawn) can time out the handshake, and the conflation makes tests that shrink request_timeout_seconds wall-clock sensitive and flaky under CI load.
Root Cause
McpClient._await_initialize() runs session.initialize() under anyio.fail_after(self._settings.request_timeout_seconds) (backend/src/octave/mcp/client.py). McpSettings has no dedicated handshake knob, so start-up budget and steady-state request budget are a single conflated value.
Reproduction
- Construct
McpSettings(request_timeout_seconds=0.5) and connect to the slow-import stdio fixture server.
- Observe
McpTimeoutError on connect under load, even though the server is healthy.
- Expected: handshake succeeds (import takes ~0.5 s of CPU but exceeds the 0.5 s budget under parallel test load); the short request timeout should only apply to steady-state calls like
tools/call.
Fix
Start-up budget != steady-state request budget. Add a dedicated knob:
McpSettings.initialize_timeout_seconds: float = 30.0 in config.py
_await_initialize() uses it instead of request_timeout_seconds (and the timeout error message references it)
- Tests keep
request_timeout_seconds=0.5 — slow_tool (5 s) deterministically exceeds it with no load sensitivity, while the handshake gets a 60x margin (30 s vs ~0.5 s import)
- Docs:
docs/DEVELOPMENT.md env-var table gains OCTAVE_MCP_INITIALIZE_TIMEOUT_SECONDS
Defaults are unchanged for all existing behavior (30 s == old effective handshake budget); zero-impact by default.
A test-only mitigation (bump timeout to 5 s + sleep 60 s) would only widen the margin, not remove the conflation — still wall-clock dependent. Separate bug from #78; own fix/ branch.
Description
The MCP
initializehandshake borrows the steady-state per-request timeout (request_timeout_seconds) as its start-up budget. A slow but healthy server (heavy module import at spawn) can time out the handshake, and the conflation makes tests that shrinkrequest_timeout_secondswall-clock sensitive and flaky under CI load.Root Cause
McpClient._await_initialize()runssession.initialize()underanyio.fail_after(self._settings.request_timeout_seconds)(backend/src/octave/mcp/client.py).McpSettingshas no dedicated handshake knob, so start-up budget and steady-state request budget are a single conflated value.Reproduction
McpSettings(request_timeout_seconds=0.5)and connect to the slow-import stdio fixture server.McpTimeoutErroron connect under load, even though the server is healthy.tools/call.Fix
Start-up budget != steady-state request budget. Add a dedicated knob:
McpSettings.initialize_timeout_seconds: float = 30.0inconfig.py_await_initialize()uses it instead ofrequest_timeout_seconds(and the timeout error message references it)request_timeout_seconds=0.5—slow_tool(5 s) deterministically exceeds it with no load sensitivity, while the handshake gets a 60x margin (30 s vs ~0.5 s import)docs/DEVELOPMENT.mdenv-var table gainsOCTAVE_MCP_INITIALIZE_TIMEOUT_SECONDSDefaults are unchanged for all existing behavior (30 s == old effective handshake budget); zero-impact by default.
A test-only mitigation (bump timeout to 5 s + sleep 60 s) would only widen the margin, not remove the conflation — still wall-clock dependent. Separate bug from #78; own
fix/branch.