MCP sessions, resources, prompts, and a chart compiler - #47
Open
prasadpamidi wants to merge 17 commits into
Open
prasadpamidi wants to merge 17 commits into
prasadpamidi wants to merge 17 commits into
Conversation
`MCPClient` ran a full `initialize` handshake on every request and tore the transport down afterwards, on the stated grounds that this "matches MCP's session-per-call model". MCP has no such model. It is a stateful session protocol — initialize, operate, shut down — and Streamable HTTP carries an `Mcp-Session-Id` across the session. What the per-call approach actually bought: - A handshake on every tool call. A turn calling three tools paid three. - Any server holding per-session state — a cache, a cursor, an auth context — seeing each call as a brand-new session. That fails quietly rather than erroring, which is the worst way to fail. `MCPSessionPool` keeps one connected client per (endpoint, credential, client identity) and evicts after three minutes idle. The credential is in the key because the `Authorization` header is baked into the transport at construction, so a refreshed OAuth token needs a new transport and must not ride the old session. It is keyed by fingerprint rather than by value so secrets stay out of dictionary keys. A pooled connection can be closed by the server, a proxy, or the OS between calls, and the caller cannot distinguish "the session went stale" from "this request is bad" — so the pool retries connection failures once, on a fresh transport. It deliberately does not retry `serverError`: that is the server *answering*, and repeating a request it already processed risks doing a non-idempotent thing twice. `makeTransport` is a factory taking `any Transport` rather than a concrete `HTTPClientTransport` value, because a retry needs a fresh transport (a disconnected one cannot be reconnected) and because it makes the pool testable. `FakeMCPTransport` is an in-process server just complete enough to answer a handshake and `tools/list`, counting every connect — the pool's whole job is deciding *when* to connect, which is invisible against a live server and untestable against none. Also fixes client identity. `clientName` defaulted to the literal "Avyra" in a shared package and no call site overrode it, so every Niora request introduced itself as Avyra to the server. It now reads the host bundle's name, which fixes both apps without touching a single call site. 387 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The client implemented `tools/*` and nothing else. Tool *results* could carry an embedded resource, but resources could not be listed or read on their own, and prompts were absent entirely — so a server's content was invisible unless a tool happened to hand it over. Adds `listResources`, `readResource`, and `listPrompts`, all draining pagination the way `listTools` already did. Descriptors are plain values with no SDK types in them, so persistence and UI never import `MCP`. `MCPResourceDescriptor.isInteractiveUI` identifies an MCP App, and requires *both* halves of the signal: a `ui://` URI and a `profile=mcp-app` MIME parameter. Checking only the MIME type would treat any server serving `text/html` as an interactive surface and hand it a message channel it never asked for; checking only the scheme would miss that the profile is what makes it an app. Prompts land now rather than later because they are the shape Discover recipes already want: a named, described entry point with typed arguments. Testing needed a seam. `MCPClient` built its own `HTTPClientTransport` from the endpoint, so pagination draining, content-block conversion and prompt-argument flattening could only be exercised against a live third-party server — which is to say never in CI. An internal initializer now accepts a transport factory, and the fake server answers `resources/list` across *two* pages specifically so a client that ignored `nextCursor` would fail the test rather than silently returning half a server's resources. 391 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Agents should be able to say "this wants a chart" and have one appear. Two constraints make that safe, and both are about what the model is *not* allowed to do. **It does not author the spec.** `ChartIntent` is an enum plus four field names — roughly thirty tokens against the several hundred a Vega-Lite spec would cost, which matters against a 4,096-token window that also has to hold the data. An invalid chart type becomes unrepresentable rather than a runtime surprise; given free text a small model eventually emits `"linechart"` or a subtly wrong encoding block. The codebase already reached this conclusion once — `BriefActionSuggestion` is "kept simple so the model never has to author an `ActionTarget`". **It does not supply the data.** The intent names a tool result and the fields within it; the host injects the rows. A model that retypes numbers eventually invents them, and an invented number wearing the authority of a chart is worse than an invented sentence. This makes charts *structurally* grounded — no gate has to check them, because the values never passed through the model. Consequences that follow from the host owning the data: - Field types are inferred, not declared. A field is quantitative only if every present value is numeric — one stray "n/a" degrades it to nominal rather than rendering a broken scale. - Date detection is deliberately narrow. A false temporal turns a category axis into a broken time axis, which is worse than a date drawn as a category, so "3 items" stays nominal. - A named field that does not exist throws, carrying the field names that *do*, since that is the likely failure and the recovery is telling the model the real names. Tooltips are always on: an interactive chart the user cannot interrogate is a picture, and that is the entire argument against Swift Charts here, where every interaction would be hand-built. Improving charts now happens in the compiler — no prompt change, no model change. That separation is the payoff for keeping the output small. 400 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A chart needs `[[String: JSONValue]]`. Tools return whatever shape their
author chose — a bare array, `{"items": [...]}`, something two levels
down, or JSON handed back as a string.
The model must not be what reshapes it. Asking a model to transcribe
rows into a chart call is asking it to invent them, and an invented
number wearing the authority of a chart is worse than an invented
sentence. So the host looks instead, and the rules are deliberately dull:
- Conventional envelope keys are tried in a fixed order, so a result
carrying both `items` and `warnings` charts the items — and charts the
same thing on every run rather than depending on dictionary ordering.
Remaining keys are searched sorted, for the same reason.
- A JSON-in-a-string result is decoded, but only at the top level.
Decoding strings found deep inside a payload would be guessing.
- A mixed array is *not* a table. Charting only its object-shaped
elements would silently drop data, which is worse than declining.
- Rows are capped, and the original count is reported rather than
swallowed — ten thousand points is an unreadable chart and a slow one,
but the user should be told what they are looking at.
The one thing this refuses to do is guess creatively. A surprising rule
here draws a chart of the wrong data, and a chart of the wrong data is
believed.
410 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
With `streaming: false` there is no long-lived listener, so `notifications/tools/list_changed` can never arrive. That is why a server changing its tools stays stale in the app until someone opens Settings and taps refresh — the client had no way to be told. Streaming is now per-server and off by default, because the existing note is right that many servers do not support the GET SSE listener. This has to degrade, not fail: a server without it keeps working exactly as before. The session key includes the streaming mode. Streaming and non-streaming are different transports, not a setting on one, and sharing a session between them would hand a caller that turned streaming on specifically to receive notifications a pooled connection that cannot deliver them. Notification handlers register *before* connect. A server may emit `tools/list_changed` immediately after initialize, and a handler attached afterwards would miss exactly the case it exists for. One implementation note worth keeping: building the `onConnect` closure inline at the call site made the type checker fail outright — "failed to produce diagnostic for expression". It is constructed up front with an explicit type instead. 411 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
`maxTools` bounds how *many* tools are sent. `toolTokenLimit` bounds
what they cost. Only the first was enforced once ranking returned
anything:
guard selected.count == requiredTools.count else {
return selected // no cost check, ever
}
The ceiling was consulted for the "everything already fits" early exit
and for filler tools, then skipped entirely on the path almost every
real turn takes. Six tools is a fine cap until the six are MCP schemas.
A connected weather server (17 tools) put **5,845 tokens into a
4,096-token window** and the turn was refused outright — with the budget
reporting itself satisfied, because six is less than six.
That is the failure mode this whole layer exists to prevent, so it is
worth being precise about what was wrong: not the budget, not the
selector, not the consumer's configuration. The assembler counted tools
and never weighed them.
Ranked tools are now taken in order until the next one does not fit, and
the loop *stops* rather than skipping ahead to a smaller candidate.
Skipping would quietly prefer cheap tools over relevant ones, which
inverts the purpose of ranking.
Required tools are exempt. They are pinned or already invoked this turn,
and dropping one breaks the conversation outright, whereas overshooting
risks a refusal that trimming everything else may still avoid.
The regression test was checked against the unfixed code rather than
assumed: it selects 6 tools costing 2,130 against a ceiling of 1,198 and
fails, then passes once trimmed. A test for a budget bug that has never
been seen to fail is not evidence of anything.
413 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
`toolsOffered` was a count. A count is enough to notice a surprise and never enough to explain one. A run asked for a weather summary with a weather server connected and sent exactly one tool — `load_skill`, which was pinned, meaning the ranker had scored *nothing* across twenty-seven candidates. From the diagnostic alone there was no way to tell whether the ranker had buried the weather tools or whether they had never been candidates at all. Those have nothing in common except the symptom, and opposite fixes: one is a ranking bug, the other is a tool that was never registered. `ContextAllocation.offeredToolNames` carries the candidate list, so the question is answerable by reading the export instead of by reasoning backwards from a number. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
…ut it" A weather question with a weather server connected sent one tool. The export said 35 offered, 1 selected — and that is where the trail ended, because two completely different failures produce that same pair of numbers: - ranking found nothing, so only the pinned tool survived; or - ranking found the right tools and the token ceiling trimmed them away. One is a retrieval problem, the other a budget problem, and the fix for either is useless against the other. Three turns and a reproduction were spent distinguishing them by hand, and the answer was still not conclusive. `ContextAllocation` now carries `rankedToolNames` — what ranking chose *before* trimming — alongside `maxTools` and `toolTokenLimit`, the two caps in force. Empty ranked names means retrieval; full ranked names with a short selection means budget. No inference required. `selectTools` returns the ranking beside the selection rather than stashing it, since `DefaultContextAssembler` is a struct and scratch state would not compile — and a tuple keeps the two facts adjacent, which is how they are read. Worth recording for the corpus: the assembler was reproduced against the exact 35-tool surface from the field export, pinned `load_skill`, `maxTools: 6`, `unrankedFillLimit: 0`, and it selected `load_skill` plus five weather tools correctly. Whatever produced the field result is not in this code path. 414 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
This morning's ceiling fix traded an overflow for a turn that cannot
act, which is the worse failure.
The field diagnostic, once it carried enough to be conclusive:
RANKED before trim: [load_skill, weather_mcp__get_weather_summary,
weather_mcp__get_forecast, ...]
SENT: [load_skill]
tool ceiling: 1198 tools sent: 118 tokens
Ranking was right — it put `get_weather_summary` second for "Get
weather summary". Every schema on that MCP server runs to roughly a
thousand tokens, over the 1,080 left after the pinned tool, so `fitting`
broke on the first candidate and returned the pinned tool alone. The
model was handed no way to fetch weather and answered from imagination,
filling a template with `[Current Date]` and `[Temperature Range]`.
The trim was correct by its own rule, and the rule was wrong. The
ceiling is a *share* of the budget — 40% by default — not the context
window. Honouring it exactly is right when it costs a fourth tool and
wrong when it costs the only one: no amount of budget discipline
redeems a request that cannot do the thing it was asked to do.
So when the share admits nothing, the top-ranked candidate is admitted
anyway. Deliberately the top-ranked one rather than the largest that
fits — preferring a cheap tool to a relevant one inverts ranking, and
this is a last resort, not a second policy.
Checked against the unfixed code rather than assumed: it reproduces the
field result exactly, `["load_skill"]` and nothing else.
Two things this leaves open, stated rather than buried. A thousand-token
tool schema is enormous, and six of them are what put 5,845 tokens into
a 4,096-token window earlier — capping or summarising oversized schemas
is the real fix and is not this one. And admitting an over-share tool
can still overflow a small window; it is strictly better than sending
nothing, and strictly worse than schemas that fit.
414 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A published MCP tool schema is routinely a thousand tokens: a paragraph of description plus a dozen documented parameters. On the 4,096-token window an edge device actually has, that is a quarter of everything for one tool the model may not even call. Six of them produced a 5,845-token request that was refused outright, and the budget fix that followed turned that into a turn with one tool and no way to answer. Budgeting cannot solve this. A share that fits six such tools does not exist. But most of those tokens are prose the model does not need in order to make the call — so shrink the definition rather than choosing between overflow and starvation. `ToolDefinitionCompactor` applies the lightest level that fits: 1. drop per-property descriptions — parameter names carry most of it 2. also shorten the tool description to its first sentence 3. also drop optional properties The ordering is the design, and the invariant is that **a call the model could make before must still validate after**. Names, types, the required list and enum values survive every level: enums are not documentation, they are the set of legal inputs, and a model guessing outside one produces a call that fails. What is given up is guidance first and optional capability last, because a tool called with worse arguments is recoverable and a tool whose arguments no longer validate is not. `AnyTool.replacingDefinition` keeps the invocation closure untouched, so a compacted definition can change what the model is *told* and never what happens when it calls. Verified in both directions, after a first attempt that passed vacuously: with schemas sized like the real ones, removing the compaction hook yields one tool and 1,390 tokens against a 1,198 ceiling — the field failure exactly — and restoring it yields several tools under budget. Still open: this shrinks what a server published, it does not make servers publish less. A tool whose *required* parameters alone exceed the budget is beyond it. 421 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Two rules composed into the worst possible outcome. The share (`toolTokenLimit`, 40% of the budget) is deliberately breakable: sending no usable tool is worse than overspending it, and a turn that cannot act is a turn that answers from imagination. But nothing bounded the result. A run selected three MCP tools worth 4,533 tokens against a 1,198 share and a 2,996 budget; `remaining` went to zero, history was discarded to make room, and the provider refused the request anyway at 4,727 of 4,096. The conversation was thrown away *and* the turn failed — strictly worse than either failure alone. `withinHardLimit` drops tools, lowest-ranked first, until the request physically fits, recomposing the prompt each time because guidance rides with the tools that survived. The newest message is reserved for: a request that cannot carry the user's turn is not a smaller request, it is a broken one. So the hierarchy is now explicit. The share is advice. The budget is not. Compaction tries to make both satisfiable before either has to give. Verified in both directions, after two vacuous attempts — the first fixture compacted small enough to fit, the second was one incompressible tool that still fit. It took a tool whose *required* surface alone exceeds the whole budget to reproduce it: 10,401 tokens against 2,996 without the bound, within budget with it. That is three tests today that looked green while asserting nothing. Reverting the fix before trusting the test is cheap; believing a green test that never failed is not. 422 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A 0.8B model found the right tool, formed the right intent, and picked
the right values — then the call failed on quotation marks:
{"days": "1", "latitude": "56.35", "longitude": "-4.11"}
→ "Invalid days: must be a finite number, received string"
Nothing upstream was wrong. The Qwen parser preserves JSON types
faithfully; the model simply quoted its numbers, which small models do.
The schema already says what each field is, so the host can reconcile
that instead of asking the model to try again — the same division the
rest of this layer runs on: the model decides *what*, the host owns the
shape.
Only lossless conversions. `"1"` becomes `1`; `"next week"` stays a
string and the server returns its own error, which is more useful than a
number this invented. Specifically not done:
- `1.7` is never truncated to an integer — that changes the request.
- `"yes"` is not read as `true`. It is a guess about intent rather than
a reading of the value, and JSON has two spellings for a reason.
- A field the schema does not describe passes through untouched; there
is no contract to enforce and inventing one is worse than nothing.
Also handles the mirror case — a bare number where the schema declares a
string — since it costs nothing and fails the same way.
422 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Compaction shrank what the assembler *counted* and not what the provider *sent*, which is worse than not compacting at all. `FoundationModelsProvider` builds its typed tools from factory closures handed to it at construction. Those closures captured their description then. Swapping a definition afterwards reaches the `AnyTool` the assembler returns and cannot reach the tool the provider registers — so the budget was costing compacted tools while the request carried full ones. Measured, from the field: a turn priced at 1,366 estimated tokens was refused at 5,362 actual — 3.9x under, with six tools trimmed to exactly the 1,198 ceiling and six full ones sent. An accurate budget with fewer tools beats an inaccurate budget with more. A refused turn sends none at all. `ToolDefinitionCompactor` stays, with its tests. It is correct and it will be useful to a provider that builds its tool list from `AssembledContext.tools` rather than from closures fixed at construction — which is the real fix, and is not a one-line change to make safely. Also documents the constraint at both ends: the provider's filter now says why it can only match on name, and `fitting` says why it does not compact. This cost a full debugging cycle to find; the next person should not have to repeat it. 429 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The estimate is the most load-bearing number in the context layer — every trim, cap and drop is arithmetic on top of it — and nothing verified it. When it was wrong the symptom appeared somewhere else entirely: a provider refusing a turn the budget had called comfortable, with the budget still reporting itself satisfied afterwards. `ContextAudit` renders the request the way a chat template will and compares that against what the counter charged. Rendering first is the point: a counter that reads `ToolDefinition` objects cannot see a schema counted once as a schema and sent again inside a description, and that double-shipping priced as prose is a real error this catches. Assertions are one-sided, because the directions are not symmetric. Over-estimating wastes budget and sends fewer tools; under-estimating overflows the window and loses the turn. What it does **not** catch, stated plainly: it renders the tools the assembler produced, so it cannot see a provider sending something else. That was the 1,366-vs-5,362 failure — definitions compacted after the provider had captured the originals — and both sides of this audit would have agreed while the request was four times larger. Closing that needs the provider's own count, which is the companion change in the app. The reference counter is a character heuristic, so this catches order-of-magnitude errors and not small ones. It takes a closure so a real tokenizer can be supplied where one exists. 434 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The eval corpus measures whether the right tool is *offered*. It has almost nothing on the opposite failure — a tool offered when none is relevant — and that is the one users notice. From the field: asked "What's the tracking number?", the ranker offered `calculator`, the model called it with `(4.45 - 4.45) / 16` — numbers scavenged out of a fasting summary earlier in the conversation — and answered `0`. Measured against a real 34-tool device surface, the pattern is exact. Everything that fails does so on one common word matching one tool: "number" appears in `calculator`, `uuid` and `unit_converter`; "fact" appears in `remember_fact`. Everything that succeeds matches on distinctive terms, usually more than one. Conversational turns — "Thanks!", "How are you?", "Explain quantum computing" — already rank nothing and now stay that way. The two failures are recorded as `XCTExpectFailure` rather than fixed here, and that is deliberate. The obvious rule — require two matched terms — breaks "remind me to buy milk", where one matched term is the whole intent. Picking a threshold against the four cases that happened to motivate it is how the stopword list came to delete the subject of every fasting query, twice. So this states the gap instead of guessing at it. The other half of the corpus is the constraint any future floor has to satisfy: four real requests that must keep ranking their tool first. 435 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Attached images were invisible in two independent ways, and each alone is enough to break a turn. **They cost nothing.** `count(message:)` measured `textContent` and the tool calls, so an image contributed zero. Everything downstream is arithmetic on that number: history windowing never dropped an image message because it looked free, and the request overflowed with the budget reporting itself satisfied. Images are now charged `tokensPerImage`, a deliberately large constant — a 1024x1024 image runs to roughly a thousand tokens on the Qwen-VL family, and nothing about that is derivable from the bytes we hold. Over-charging costs a few tools; under-charging costs the turn. **A photo with no caption got no tools at all.** The ranking query is the latest user text, which is empty for an image-only turn. The selector returns nothing for an empty query, `unrankedFillLimit: 0` reads that as a decision, and the model received an image with no way to act on it. Those are different situations. "Nothing matched" is a ranker verdict; an empty query means the ranker was never given a question. The empty case now falls through to what fits, which is the same reading the definition-only provider path already takes. A captioned image is unaffected — it still ranks on its caption, which has its own test, because a fallback that swallowed the normal path would be worse than the bug. Both tests were vacuous on the first attempt and were rewritten until they failed against the unfixed code. The trap each time: a surface smaller than `maxTools` short-circuits the assembler and returns everything unranked, so a selection test with four tools and a cap of six proves nothing. That is the fourth time today. 441 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Asked "What's the tracking number" over a receipt photo, a model answered from the caption alone and invented coordinates in the Bay of Biscay to call a weather tool with. The diagnostic could not say whether the image had reached it: the allocation described tokens, tools and caps, and was silent on the one thing that mattered. `ContextAllocation.imageCount` closes that. Zero with an attachment on screen localises the loss to assembly; non-zero moves it downstream to the provider or the model. Without it the question needs a code read and a guess, which is how the last four of these went. 441 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Foundation work from
docs/plans/2026-08-09-mcp-brief-ui-charts.md(in the Avyra repo). Three commits, each independently useful.1. One session per server, not one per call
MCPClientran a fullinitializehandshake on every request, on the stated grounds that this "matches MCP's session-per-call model". MCP has no such model — it's a stateful session protocol, and Streamable HTTP carries anMcp-Session-Idacross the session.What that cost: a handshake per tool call (a turn calling three tools paid three), and any server holding per-session state seeing each call as a brand-new session — failing quietly rather than erroring.
MCPSessionPoolkeeps one connected client per (endpoint, credential, identity), evicted after 3 minutes idle. Credential is in the key because theAuthorizationheader is baked into the transport at construction, so a refreshed OAuth token can't ride the old session — fingerprinted, so secrets stay out of dictionary keys.Retries connection failures once on a fresh transport; never retries
serverError, because that's the server answering and repeating it risks doing a non-idempotent thing twice.Also fixes client identity:
clientNamedefaulted to the literal"Avyra"in a shared package with no call site overriding it, so every Niora request introduced itself as Avyra. Now reads the host bundle name — both apps fixed, zero call sites touched.2. Resources and prompts
Only
tools/*was implemented. AddslistResources,readResource,listPrompts, all draining pagination.isInteractiveUIidentifies an MCP App and requires both signals — aui://URI and aprofile=mcp-appMIME parameter. MIME alone would treat any server servingtext/htmlas interactive and hand it a message channel it never asked for.3. ChartIntent → Vega-Lite
The model emits an enum plus four field names (~30 tokens), never a spec and never the data. It names a tool result; the host injects the rows. A model that retypes numbers eventually invents them, and an invented number wearing the authority of a chart is worse than an invented sentence — so charts are structurally grounded rather than gated.
Field types are inferred from values: quantitative only if every value is numeric (one stray
"n/a"degrades to nominal rather than drawing a broken scale), and date detection is narrow enough that"3 items"stays a category.Testing
Needed a seam.
MCPClientbuilt its own transport, so pagination and content mapping could only be tested against a live third-party server. An internal initializer now takes a transport factory, andFakeMCPTransportis an in-process server that counts every connect — the pool's whole job is deciding when to connect, invisible against a live server and untestable against none.The assertion that matters: three calls, one handshake. And
resources/listis faked across two pages so a client ignoringnextCursorfails rather than silently returning half a server's resources.400 tests pass.
🤖 Generated with Claude Code
https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L