Skip to content

MCP sessions, resources, prompts, and a chart compiler - #47

Open
prasadpamidi wants to merge 17 commits into
mainfrom
mcp/session-persistence
Open

prasadpamidi wants to merge 17 commits into
mainfrom
mcp/session-persistence

Conversation

@prasadpamidi

Copy link
Copy Markdown
Owner

Foundation work from docs/plans/2026-08-09-mcp-brief-ui-charts.md (in the Avyra repo). Three commits, each independently useful.

1. One session per server, not one per call

MCPClient ran a full initialize handshake on every request, on the stated grounds that this "matches MCP's session-per-call model". MCP has no such model — it's a stateful session protocol, and Streamable HTTP carries an Mcp-Session-Id across the session.

What that cost: a handshake per tool call (a turn calling three tools paid three), and any server holding per-session state seeing each call as a brand-new session — failing quietly rather than erroring.

MCPSessionPool keeps one connected client per (endpoint, credential, identity), evicted after 3 minutes idle. Credential is in the key because the Authorization header is baked into the transport at construction, so a refreshed OAuth token can't ride the old session — fingerprinted, so secrets stay out of dictionary keys.

Retries connection failures once on a fresh transport; never retries serverError, because that's the server answering and repeating it risks doing a non-idempotent thing twice.

Also fixes client identity: clientName defaulted to the literal "Avyra" in a shared package with no call site overriding it, so every Niora request introduced itself as Avyra. Now reads the host bundle name — both apps fixed, zero call sites touched.

2. Resources and prompts

Only tools/* was implemented. Adds listResources, readResource, listPrompts, all draining pagination.

isInteractiveUI identifies an MCP App and requires both signals — a ui:// URI and a profile=mcp-app MIME parameter. MIME alone would treat any server serving text/html as interactive and hand it a message channel it never asked for.

3. ChartIntent → Vega-Lite

The model emits an enum plus four field names (~30 tokens), never a spec and never the data. It names a tool result; the host injects the rows. A model that retypes numbers eventually invents them, and an invented number wearing the authority of a chart is worse than an invented sentence — so charts are structurally grounded rather than gated.

Field types are inferred from values: quantitative only if every value is numeric (one stray "n/a" degrades to nominal rather than drawing a broken scale), and date detection is narrow enough that "3 items" stays a category.

Testing

Needed a seam. MCPClient built its own transport, so pagination and content mapping could only be tested against a live third-party server. An internal initializer now takes a transport factory, and FakeMCPTransport is an in-process server that counts every connect — the pool's whole job is deciding when to connect, invisible against a live server and untestable against none.

The assertion that matters: three calls, one handshake. And resources/list is faked across two pages so a client ignoring nextCursor fails rather than silently returning half a server's resources.

400 tests pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L

prasadpamidi and others added 17 commits August 9, 2026 11:09
`MCPClient` ran a full `initialize` handshake on every request and tore
the transport down afterwards, on the stated grounds that this "matches
MCP's session-per-call model".

MCP has no such model. It is a stateful session protocol — initialize,
operate, shut down — and Streamable HTTP carries an `Mcp-Session-Id`
across the session. What the per-call approach actually bought:

- A handshake on every tool call. A turn calling three tools paid three.
- Any server holding per-session state — a cache, a cursor, an auth
  context — seeing each call as a brand-new session. That fails
  quietly rather than erroring, which is the worst way to fail.

`MCPSessionPool` keeps one connected client per (endpoint, credential,
client identity) and evicts after three minutes idle. The credential is
in the key because the `Authorization` header is baked into the
transport at construction, so a refreshed OAuth token needs a new
transport and must not ride the old session. It is keyed by fingerprint
rather than by value so secrets stay out of dictionary keys.

A pooled connection can be closed by the server, a proxy, or the OS
between calls, and the caller cannot distinguish "the session went
stale" from "this request is bad" — so the pool retries connection
failures once, on a fresh transport. It deliberately does not retry
`serverError`: that is the server *answering*, and repeating a request
it already processed risks doing a non-idempotent thing twice.

`makeTransport` is a factory taking `any Transport` rather than a
concrete `HTTPClientTransport` value, because a retry needs a fresh
transport (a disconnected one cannot be reconnected) and because it
makes the pool testable. `FakeMCPTransport` is an in-process server
just complete enough to answer a handshake and `tools/list`, counting
every connect — the pool's whole job is deciding *when* to connect,
which is invisible against a live server and untestable against none.

Also fixes client identity. `clientName` defaulted to the literal
"Avyra" in a shared package and no call site overrode it, so every
Niora request introduced itself as Avyra to the server. It now reads
the host bundle's name, which fixes both apps without touching a
single call site.

387 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The client implemented `tools/*` and nothing else. Tool *results* could
carry an embedded resource, but resources could not be listed or read on
their own, and prompts were absent entirely — so a server's content was
invisible unless a tool happened to hand it over.

Adds `listResources`, `readResource`, and `listPrompts`, all draining
pagination the way `listTools` already did. Descriptors are plain values
with no SDK types in them, so persistence and UI never import `MCP`.

`MCPResourceDescriptor.isInteractiveUI` identifies an MCP App, and
requires *both* halves of the signal: a `ui://` URI and a
`profile=mcp-app` MIME parameter. Checking only the MIME type would
treat any server serving `text/html` as an interactive surface and hand
it a message channel it never asked for; checking only the scheme would
miss that the profile is what makes it an app.

Prompts land now rather than later because they are the shape Discover
recipes already want: a named, described entry point with typed
arguments.

Testing needed a seam. `MCPClient` built its own `HTTPClientTransport`
from the endpoint, so pagination draining, content-block conversion and
prompt-argument flattening could only be exercised against a live
third-party server — which is to say never in CI. An internal
initializer now accepts a transport factory, and the fake server
answers `resources/list` across *two* pages specifically so a client
that ignored `nextCursor` would fail the test rather than silently
returning half a server's resources.

391 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Agents should be able to say "this wants a chart" and have one appear.
Two constraints make that safe, and both are about what the model is
*not* allowed to do.

**It does not author the spec.** `ChartIntent` is an enum plus four
field names — roughly thirty tokens against the several hundred a
Vega-Lite spec would cost, which matters against a 4,096-token window
that also has to hold the data. An invalid chart type becomes
unrepresentable rather than a runtime surprise; given free text a small
model eventually emits `"linechart"` or a subtly wrong encoding block.
The codebase already reached this conclusion once — `BriefActionSuggestion`
is "kept simple so the model never has to author an `ActionTarget`".

**It does not supply the data.** The intent names a tool result and the
fields within it; the host injects the rows. A model that retypes
numbers eventually invents them, and an invented number wearing the
authority of a chart is worse than an invented sentence. This makes
charts *structurally* grounded — no gate has to check them, because the
values never passed through the model.

Consequences that follow from the host owning the data:

- Field types are inferred, not declared. A field is quantitative only
  if every present value is numeric — one stray "n/a" degrades it to
  nominal rather than rendering a broken scale.
- Date detection is deliberately narrow. A false temporal turns a
  category axis into a broken time axis, which is worse than a date
  drawn as a category, so "3 items" stays nominal.
- A named field that does not exist throws, carrying the field names
  that *do*, since that is the likely failure and the recovery is
  telling the model the real names.

Tooltips are always on: an interactive chart the user cannot interrogate
is a picture, and that is the entire argument against Swift Charts here,
where every interaction would be hand-built.

Improving charts now happens in the compiler — no prompt change, no
model change. That separation is the payoff for keeping the output
small.

400 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A chart needs `[[String: JSONValue]]`. Tools return whatever shape their
author chose — a bare array, `{"items": [...]}`, something two levels
down, or JSON handed back as a string.

The model must not be what reshapes it. Asking a model to transcribe
rows into a chart call is asking it to invent them, and an invented
number wearing the authority of a chart is worse than an invented
sentence. So the host looks instead, and the rules are deliberately dull:

- Conventional envelope keys are tried in a fixed order, so a result
  carrying both `items` and `warnings` charts the items — and charts the
  same thing on every run rather than depending on dictionary ordering.
  Remaining keys are searched sorted, for the same reason.
- A JSON-in-a-string result is decoded, but only at the top level.
  Decoding strings found deep inside a payload would be guessing.
- A mixed array is *not* a table. Charting only its object-shaped
  elements would silently drop data, which is worse than declining.
- Rows are capped, and the original count is reported rather than
  swallowed — ten thousand points is an unreadable chart and a slow one,
  but the user should be told what they are looking at.

The one thing this refuses to do is guess creatively. A surprising rule
here draws a chart of the wrong data, and a chart of the wrong data is
believed.

410 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
With `streaming: false` there is no long-lived listener, so
`notifications/tools/list_changed` can never arrive. That is why a
server changing its tools stays stale in the app until someone opens
Settings and taps refresh — the client had no way to be told.

Streaming is now per-server and off by default, because the existing
note is right that many servers do not support the GET SSE listener.
This has to degrade, not fail: a server without it keeps working exactly
as before.

The session key includes the streaming mode. Streaming and
non-streaming are different transports, not a setting on one, and
sharing a session between them would hand a caller that turned
streaming on specifically to receive notifications a pooled connection
that cannot deliver them.

Notification handlers register *before* connect. A server may emit
`tools/list_changed` immediately after initialize, and a handler
attached afterwards would miss exactly the case it exists for.

One implementation note worth keeping: building the `onConnect` closure
inline at the call site made the type checker fail outright — "failed to
produce diagnostic for expression". It is constructed up front with an
explicit type instead.

411 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
`maxTools` bounds how *many* tools are sent. `toolTokenLimit` bounds
what they cost. Only the first was enforced once ranking returned
anything:

    guard selected.count == requiredTools.count else {
        return selected        // no cost check, ever
    }

The ceiling was consulted for the "everything already fits" early exit
and for filler tools, then skipped entirely on the path almost every
real turn takes. Six tools is a fine cap until the six are MCP schemas.
A connected weather server (17 tools) put **5,845 tokens into a
4,096-token window** and the turn was refused outright — with the budget
reporting itself satisfied, because six is less than six.

That is the failure mode this whole layer exists to prevent, so it is
worth being precise about what was wrong: not the budget, not the
selector, not the consumer's configuration. The assembler counted tools
and never weighed them.

Ranked tools are now taken in order until the next one does not fit, and
the loop *stops* rather than skipping ahead to a smaller candidate.
Skipping would quietly prefer cheap tools over relevant ones, which
inverts the purpose of ranking.

Required tools are exempt. They are pinned or already invoked this turn,
and dropping one breaks the conversation outright, whereas overshooting
risks a refusal that trimming everything else may still avoid.

The regression test was checked against the unfixed code rather than
assumed: it selects 6 tools costing 2,130 against a ceiling of 1,198 and
fails, then passes once trimmed. A test for a budget bug that has never
been seen to fail is not evidence of anything.

413 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
`toolsOffered` was a count. A count is enough to notice a surprise and
never enough to explain one.

A run asked for a weather summary with a weather server connected and
sent exactly one tool — `load_skill`, which was pinned, meaning the
ranker had scored *nothing* across twenty-seven candidates. From the
diagnostic alone there was no way to tell whether the ranker had buried
the weather tools or whether they had never been candidates at all.
Those have nothing in common except the symptom, and opposite fixes:
one is a ranking bug, the other is a tool that was never registered.

`ContextAllocation.offeredToolNames` carries the candidate list, so the
question is answerable by reading the export instead of by reasoning
backwards from a number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
…ut it"

A weather question with a weather server connected sent one tool. The
export said 35 offered, 1 selected — and that is where the trail ended,
because two completely different failures produce that same pair of
numbers:

- ranking found nothing, so only the pinned tool survived; or
- ranking found the right tools and the token ceiling trimmed them away.

One is a retrieval problem, the other a budget problem, and the fix for
either is useless against the other. Three turns and a reproduction were
spent distinguishing them by hand, and the answer was still not
conclusive.

`ContextAllocation` now carries `rankedToolNames` — what ranking chose
*before* trimming — alongside `maxTools` and `toolTokenLimit`, the two
caps in force. Empty ranked names means retrieval; full ranked names
with a short selection means budget. No inference required.

`selectTools` returns the ranking beside the selection rather than
stashing it, since `DefaultContextAssembler` is a struct and scratch
state would not compile — and a tuple keeps the two facts adjacent,
which is how they are read.

Worth recording for the corpus: the assembler was reproduced against
the exact 35-tool surface from the field export, pinned `load_skill`,
`maxTools: 6`, `unrankedFillLimit: 0`, and it selected `load_skill`
plus five weather tools correctly. Whatever produced the field result
is not in this code path.

414 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
This morning's ceiling fix traded an overflow for a turn that cannot
act, which is the worse failure.

The field diagnostic, once it carried enough to be conclusive:

    RANKED before trim: [load_skill, weather_mcp__get_weather_summary,
                         weather_mcp__get_forecast, ...]
    SENT:               [load_skill]
    tool ceiling: 1198   tools sent: 118 tokens

Ranking was right — it put `get_weather_summary` second for "Get
weather summary". Every schema on that MCP server runs to roughly a
thousand tokens, over the 1,080 left after the pinned tool, so `fitting`
broke on the first candidate and returned the pinned tool alone. The
model was handed no way to fetch weather and answered from imagination,
filling a template with `[Current Date]` and `[Temperature Range]`.

The trim was correct by its own rule, and the rule was wrong. The
ceiling is a *share* of the budget — 40% by default — not the context
window. Honouring it exactly is right when it costs a fourth tool and
wrong when it costs the only one: no amount of budget discipline
redeems a request that cannot do the thing it was asked to do.

So when the share admits nothing, the top-ranked candidate is admitted
anyway. Deliberately the top-ranked one rather than the largest that
fits — preferring a cheap tool to a relevant one inverts ranking, and
this is a last resort, not a second policy.

Checked against the unfixed code rather than assumed: it reproduces the
field result exactly, `["load_skill"]` and nothing else.

Two things this leaves open, stated rather than buried. A thousand-token
tool schema is enormous, and six of them are what put 5,845 tokens into
a 4,096-token window earlier — capping or summarising oversized schemas
is the real fix and is not this one. And admitting an over-share tool
can still overflow a small window; it is strictly better than sending
nothing, and strictly worse than schemas that fit.

414 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A published MCP tool schema is routinely a thousand tokens: a paragraph
of description plus a dozen documented parameters. On the 4,096-token
window an edge device actually has, that is a quarter of everything for
one tool the model may not even call. Six of them produced a 5,845-token
request that was refused outright, and the budget fix that followed
turned that into a turn with one tool and no way to answer.

Budgeting cannot solve this. A share that fits six such tools does not
exist. But most of those tokens are prose the model does not need in
order to make the call — so shrink the definition rather than choosing
between overflow and starvation.

`ToolDefinitionCompactor` applies the lightest level that fits:

  1. drop per-property descriptions — parameter names carry most of it
  2. also shorten the tool description to its first sentence
  3. also drop optional properties

The ordering is the design, and the invariant is that **a call the
model could make before must still validate after**. Names, types, the
required list and enum values survive every level: enums are not
documentation, they are the set of legal inputs, and a model guessing
outside one produces a call that fails. What is given up is guidance
first and optional capability last, because a tool called with worse
arguments is recoverable and a tool whose arguments no longer validate
is not.

`AnyTool.replacingDefinition` keeps the invocation closure untouched, so
a compacted definition can change what the model is *told* and never
what happens when it calls.

Verified in both directions, after a first attempt that passed
vacuously: with schemas sized like the real ones, removing the
compaction hook yields one tool and 1,390 tokens against a 1,198
ceiling — the field failure exactly — and restoring it yields several
tools under budget.

Still open: this shrinks what a server published, it does not make
servers publish less. A tool whose *required* parameters alone exceed
the budget is beyond it.

421 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Two rules composed into the worst possible outcome.

The share (`toolTokenLimit`, 40% of the budget) is deliberately
breakable: sending no usable tool is worse than overspending it, and a
turn that cannot act is a turn that answers from imagination.

But nothing bounded the result. A run selected three MCP tools worth
4,533 tokens against a 1,198 share and a 2,996 budget; `remaining` went
to zero, history was discarded to make room, and the provider refused
the request anyway at 4,727 of 4,096. The conversation was thrown away
*and* the turn failed — strictly worse than either failure alone.

`withinHardLimit` drops tools, lowest-ranked first, until the request
physically fits, recomposing the prompt each time because guidance
rides with the tools that survived. The newest message is reserved
for: a request that cannot carry the user's turn is not a smaller
request, it is a broken one.

So the hierarchy is now explicit. The share is advice. The budget is
not. Compaction tries to make both satisfiable before either has to
give.

Verified in both directions, after two vacuous attempts — the first
fixture compacted small enough to fit, the second was one incompressible
tool that still fit. It took a tool whose *required* surface alone
exceeds the whole budget to reproduce it: 10,401 tokens against 2,996
without the bound, within budget with it.

That is three tests today that looked green while asserting nothing.
Reverting the fix before trusting the test is cheap; believing a green
test that never failed is not.

422 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
A 0.8B model found the right tool, formed the right intent, and picked
the right values — then the call failed on quotation marks:

    {"days": "1", "latitude": "56.35", "longitude": "-4.11"}
    → "Invalid days: must be a finite number, received string"

Nothing upstream was wrong. The Qwen parser preserves JSON types
faithfully; the model simply quoted its numbers, which small models do.
The schema already says what each field is, so the host can reconcile
that instead of asking the model to try again — the same division the
rest of this layer runs on: the model decides *what*, the host owns the
shape.

Only lossless conversions. `"1"` becomes `1`; `"next week"` stays a
string and the server returns its own error, which is more useful than a
number this invented. Specifically not done:

- `1.7` is never truncated to an integer — that changes the request.
- `"yes"` is not read as `true`. It is a guess about intent rather than
  a reading of the value, and JSON has two spellings for a reason.
- A field the schema does not describe passes through untouched; there
  is no contract to enforce and inventing one is worse than nothing.

Also handles the mirror case — a bare number where the schema declares a
string — since it costs nothing and fails the same way.

422 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Compaction shrank what the assembler *counted* and not what the
provider *sent*, which is worse than not compacting at all.

`FoundationModelsProvider` builds its typed tools from factory closures
handed to it at construction. Those closures captured their description
then. Swapping a definition afterwards reaches the `AnyTool` the
assembler returns and cannot reach the tool the provider registers — so
the budget was costing compacted tools while the request carried full
ones.

Measured, from the field: a turn priced at 1,366 estimated tokens was
refused at 5,362 actual — 3.9x under, with six tools trimmed to exactly
the 1,198 ceiling and six full ones sent.

An accurate budget with fewer tools beats an inaccurate budget with
more. A refused turn sends none at all.

`ToolDefinitionCompactor` stays, with its tests. It is correct and it
will be useful to a provider that builds its tool list from
`AssembledContext.tools` rather than from closures fixed at
construction — which is the real fix, and is not a one-line change to
make safely.

Also documents the constraint at both ends: the provider's filter now
says why it can only match on name, and `fitting` says why it does not
compact. This cost a full debugging cycle to find; the next person
should not have to repeat it.

429 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The estimate is the most load-bearing number in the context layer —
every trim, cap and drop is arithmetic on top of it — and nothing
verified it. When it was wrong the symptom appeared somewhere else
entirely: a provider refusing a turn the budget had called comfortable,
with the budget still reporting itself satisfied afterwards.

`ContextAudit` renders the request the way a chat template will and
compares that against what the counter charged. Rendering first is the
point: a counter that reads `ToolDefinition` objects cannot see a schema
counted once as a schema and sent again inside a description, and that
double-shipping priced as prose is a real error this catches.

Assertions are one-sided, because the directions are not symmetric.
Over-estimating wastes budget and sends fewer tools; under-estimating
overflows the window and loses the turn.

What it does **not** catch, stated plainly: it renders the tools the
assembler produced, so it cannot see a provider sending something else.
That was the 1,366-vs-5,362 failure — definitions compacted after the
provider had captured the originals — and both sides of this audit would
have agreed while the request was four times larger. Closing that needs
the provider's own count, which is the companion change in the app.

The reference counter is a character heuristic, so this catches
order-of-magnitude errors and not small ones. It takes a closure so a
real tokenizer can be supplied where one exists.

434 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
The eval corpus measures whether the right tool is *offered*. It has
almost nothing on the opposite failure — a tool offered when none is
relevant — and that is the one users notice.

From the field: asked "What's the tracking number?", the ranker offered
`calculator`, the model called it with `(4.45 - 4.45) / 16` — numbers
scavenged out of a fasting summary earlier in the conversation — and
answered `0`.

Measured against a real 34-tool device surface, the pattern is exact.
Everything that fails does so on one common word matching one tool:
"number" appears in `calculator`, `uuid` and `unit_converter`; "fact"
appears in `remember_fact`. Everything that succeeds matches on
distinctive terms, usually more than one. Conversational turns —
"Thanks!", "How are you?", "Explain quantum computing" — already rank
nothing and now stay that way.

The two failures are recorded as `XCTExpectFailure` rather than fixed
here, and that is deliberate. The obvious rule — require two matched
terms — breaks "remind me to buy milk", where one matched term is the
whole intent. Picking a threshold against the four cases that happened
to motivate it is how the stopword list came to delete the subject of
every fasting query, twice.

So this states the gap instead of guessing at it. The other half of the
corpus is the constraint any future floor has to satisfy: four real
requests that must keep ranking their tool first.

435 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Attached images were invisible in two independent ways, and each alone
is enough to break a turn.

**They cost nothing.** `count(message:)` measured `textContent` and the
tool calls, so an image contributed zero. Everything downstream is
arithmetic on that number: history windowing never dropped an image
message because it looked free, and the request overflowed with the
budget reporting itself satisfied. Images are now charged
`tokensPerImage`, a deliberately large constant — a 1024x1024 image runs
to roughly a thousand tokens on the Qwen-VL family, and nothing about
that is derivable from the bytes we hold. Over-charging costs a few
tools; under-charging costs the turn.

**A photo with no caption got no tools at all.** The ranking query is
the latest user text, which is empty for an image-only turn. The
selector returns nothing for an empty query, `unrankedFillLimit: 0`
reads that as a decision, and the model received an image with no way to
act on it.

Those are different situations. "Nothing matched" is a ranker verdict;
an empty query means the ranker was never given a question. The empty
case now falls through to what fits, which is the same reading the
definition-only provider path already takes.

A captioned image is unaffected — it still ranks on its caption, which
has its own test, because a fallback that swallowed the normal path
would be worse than the bug.

Both tests were vacuous on the first attempt and were rewritten until
they failed against the unfixed code. The trap each time: a surface
smaller than `maxTools` short-circuits the assembler and returns
everything unranked, so a selection test with four tools and a cap of
six proves nothing. That is the fourth time today.

441 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Asked "What's the tracking number" over a receipt photo, a model
answered from the caption alone and invented coordinates in the Bay of
Biscay to call a weather tool with. The diagnostic could not say whether
the image had reached it: the allocation described tokens, tools and
caps, and was silent on the one thing that mattered.

`ContextAllocation.imageCount` closes that. Zero with an attachment on
screen localises the loss to assembly; non-zero moves it downstream to
the provider or the model. Without it the question needs a code read
and a guess, which is how the last four of these went.

441 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UQx5NRp1nBm1Y7sHRJu5L
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant