From 75295f56862de8c3ce8676055916d52d2e618b26 Mon Sep 17 00:00:00 2001 From: Douglas Ezra Morrison Date: Sun, 9 Aug 2026 16:23:43 -0700 Subject: [PATCH 1/3] docs: correct offline-agent guidance from measurements on a 24 GB M2 Three corrections and two additions, all measured on the machine rather than recalled, while setting it up for offline agentic work. Corrections: - The "2048-token context window" claim is stale. Ollama now picks the default from detected memory (4k/32k/256k), per `ollama serve --help`. A 24 GB Apple-silicon machine gets 4096, not 32k, because only ~75% of unified memory is GPU-addressable. Adds `ollama ps` as the way to check what you actually got, and a Modelfile route for clients that cannot send num_ctx themselves. - qwen2.5-coder does not emit tool calls (0/4 attempts, finish_reason "stop"), so it cannot drive an agentic loop, though it remains a strong completion model. granite4:7b-a1b-h succeeded 3/3 plus a multi-turn round trip at a third of the size. `ollama show` capability metadata is wrong in both directions, so only a real request settles it. - The hardware-tier table overstated 24 GB of unified memory. The 30-32B tier is unreachable there: qwen3-coder's smallest tag is 19 GB against a ~18 GB ceiling, and raising a 14B from 4k to 32k context cost ~5.5 GB of KV cache. Additions: - Ollama serves an Anthropic-compatible /v1/messages endpoint, so Claude Code runs against a local model with no proxy. - Harness overhead bounds what a small model can do, with the tool surface added as a sixth guardrail. Also documents edit_format: whole for small models: with `diff`, granite4 fixed the target function and silently deleted an adjacent one in 3/3 runs, committing a message that mentioned only the fix. Closes #56 --- chapters/ai-tools/running-agents-offline.qmd | 284 +++++++++++++++++- .../ai-tools/small-local-models-agentic.qmd | 89 +++++- inst/WORDLIST | 3 + 3 files changed, 366 insertions(+), 10 deletions(-) diff --git a/chapters/ai-tools/running-agents-offline.qmd b/chapters/ai-tools/running-agents-offline.qmd index 4b03a1b..10fef0a 100644 --- a/chapters/ai-tools/running-agents-offline.qmd +++ b/chapters/ai-tools/running-agents-offline.qmd @@ -68,6 +68,46 @@ for current requirements. As a rough guide, smaller (7B) models run on consumer GPUs with around 8 GB of VRAM, while larger (32B and 70B) models need substantially more and may not fit on a single GPU. +::: {.callout-warning} +#### A good completion model is not necessarily a usable agent + +The models above are strong at *writing code* when you ask them to. +That is a different skill from **tool calling** --- +emitting a well-formed request to read a file or run a command, +and then using the result. +Inline completion and chat need only the first. +Anything autonomous needs the second, +because an agent that cannot call a tool cannot read your repository at all. + +The two come apart in practice. +Tested against a single-function tool schema on a 24 GB M2, +`qwen2.5-coder:14b` returned `finish_reason: stop` and no tool call +on four attempts out of four, +while `granite4:7b-a1b-h` returned a well-formed call three times out of three +and completed a full multi-turn round trip --- +despite being a third of the size. + +Do not settle this from the model's advertised capabilities, +which are wrong in both directions. +On the same machine `ollama show qwen2.5-coder:14b` lists `tools` +among its capabilities and the model does not call them, +while `ollama show gemma4:12b` lists no tools and the model calls them correctly. +Test it yourself with one request before building a loop on it: + +```bash +curl -s http://localhost:11434/v1/chat/completions -H 'content-type: application/json' -d '{ + "model": "granite4:7b-a1b-h", "stream": false, + "messages": [{"role": "user", "content": "What is in src/main.py? Use the tool."}], + "tools": [{"type": "function", "function": {"name": "read_file", + "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, + "required": ["path"]}}}]}' | grep -o '"tool_calls".*' | head -c 200 +``` + +Output containing `tool_calls` means the model is usable as an agent. +Silence means it is a completion model, +whatever its model card says. +::: + **Start the Ollama server:** ```bash @@ -134,19 +174,67 @@ aider --model ollama_chat/qwen2.5-coder:7b ::: {.callout-important} #### Raise the Ollama context window -By default Ollama uses a 2048-token context window, +Ollama's default context window is small, which silently truncates your code and makes the model look far less capable than it is. This is the single most common mistake when pairing `aider` with Ollama. -Raise it with a model-settings file at `~/.aider.model.settings.yml`: -```yaml -- name: ollama_chat/qwen2.5-coder:7b - extra_params: - num_ctx: 8192 +The default is not a fixed number. +Ollama picks it from the memory it detects, +as its own `ollama serve --help` states: + +``` +OLLAMA_CONTEXT_LENGTH Context length to use unless otherwise specified + (default: 4k/32k/256k based on VRAM) +``` + +Do not assume you landed in a generous tier. +A 24 GB Apple-silicon machine gets **4096 tokens**, +not the 32k its total memory suggests, +because only about 75% of unified memory is addressable by the GPU +and the tier boundary sits above that share. +Check what you actually got rather than inferring it --- +`ollama ps` prints the context of each loaded model: + +```bash +ollama ps +# NAME SIZE PROCESSOR CONTEXT +# qwen2.5-coder:14b 9.5 GB 100% GPU 4096 ``` -Larger values (16384, 32768) handle bigger files at the cost of more memory; -pick the largest your machine can comfortably hold. +There are three ways to raise it, +and they differ in which clients they reach: + +- **`OLLAMA_CONTEXT_LENGTH`** on the server, which sets the default for everything. + Note that the macOS menu-bar app starts the server with its own environment, + so exporting the variable in your shell does not reach it; + this route applies when you run `ollama serve` yourself. +- **A `num_ctx` parameter sent per request**, + which is what `aider` does through `~/.aider.model.settings.yml`: + + ```yaml + - name: ollama_chat/qwen2.5-coder:7b + extra_params: + num_ctx: 32768 + ``` + +- **A Modelfile that bakes the context into a derived model**, + which is the only one of the three that reaches clients + that cannot send `num_ctx` themselves: + + ```bash + printf 'FROM granite4:7b-a1b-h\nPARAMETER num_ctx 32768\n' > Modelfile + ollama create granite4-32k -f Modelfile + ``` + + The derived model shares weight blobs with its base, + so it costs no extra disk. + +Context is not free. +Raising a 14B model from 4k to 32k on a 24 GB M2 took its resident size +from 9.5 GB to 15 GB, +about 5.5 GB of key-value cache, +so pick the largest value that still leaves the weights and the cache in GPU memory +and confirm with `ollama ps` that `PROCESSOR` still reads `100% GPU`. ::: To avoid passing flags every time, @@ -174,7 +262,128 @@ because the two models take turns and their weights are swapped in and out of me Reserve it for genuinely tricky changes; for small edits, a single model is faster. -#### Falling Back Between Cloud and Local Automatically +::: {.callout-important} +#### Set `edit_format: whole` for a small model + +`aider` asks the model to express an edit either as a +SEARCH/REPLACE block (`diff`, the default for most models) +or by rewriting the file (`whole`). +Producing an exact SEARCH/REPLACE block is a demanding format, +and small models are unreliable at it. + +Measured on a 24 GB M2 with `granite4:7b-a1b-h` at 32k context, +same prompt each time: + +| File | `edit_format: diff` | `edit_format: whole` | +|---|---|---| +| One function | 3 of 3 correct | 3 of 3 correct | +| Two functions, one to be left alone | 0 of 3 correct | 3 of 3 correct | + +The two-function failure is worth dwelling on, +because it is not the failure you would expect. +The model did not refuse the edit or produce a broken file. +In all three `diff` runs it fixed the target function correctly +**and silently deleted the other one**, +then committed with a message naming only the intended fix. +Nothing in the commit message, +the exit status, +or the model's own summary mentioned the deletion. + +That is the hazard of an unattended loop in its most concrete form: +a step that reports success while destroying work, +leaving a wrong state as the premise for every step after it. +It is also why committing after every step matters --- +`git` held the original, +so the damage was one `git revert` away rather than lost. + +Set the format per model: + +```yaml +- name: ollama_chat/granite4-32k + edit_format: whole +``` + +`whole` costs more tokens per edit, +since the model rewrites the whole file, +which is a real cost on large files. +Larger models generally handle `diff` correctly; +re-measure rather than assuming either way. +::: + +#### Connecting Claude Code to Ollama + +Ollama also serves an **Anthropic-compatible** endpoint at `/v1/messages`, +alongside the OpenAI-compatible one used above. +Any client that speaks the Anthropic API can therefore be pointed at a local model, +including [Claude Code](https://claude.com/claude-code) itself, +with no proxy in between. +Check that the endpoint answers before wiring anything to it: + +```bash +curl -s http://localhost:11434/v1/messages -H 'content-type: application/json' \ + -d '{"model":"granite4-32k","max_tokens":50, + "messages":[{"role":"user","content":"Say OK only."}]}' +``` + +Point Claude Code at it with three environment variables: + +```bash +export ANTHROPIC_BASE_URL=http://localhost:11434 +export ANTHROPIC_AUTH_TOKEN=ollama # any non-empty value; Ollama ignores it +export ANTHROPIC_API_KEY="" # ensure no real key is sent to localhost +claude --model granite4-32k +``` + +Set these in a wrapper script rather than in your shell profile, +so that plain `claude` keeps using the cloud +and a separate command uses the local model. + +::: {.callout-warning} +#### A heavyweight harness can defeat a small model + +This works, +but expect worse results than the same model gives through `aider`, +and understand why before blaming the model. + +A harness spends context before your task does. +Claude Code sends a long system prompt and a schema for every tool it exposes, +and any Model Context Protocol (MCP) servers you have configured +add their own schemas on top. +Against a 32k local context that overhead is a large fraction of the budget, +and a small model handles it poorly. + +Observed on a 24 GB M2 with `granite4:7b-a1b-h`, +asking only that it read one file and comment on one function: + +- With a GitHub MCP server loaded, the model ignored the question + and produced a paragraph about missing credentials. +- With MCP disabled, it invented a task list whose contents were copied + from the description of a tool it had been shown. +- With the tool surface cut to `Read`, `Grep`, and `Glob`, it still ignored + the question and asked what it should work on. + +The same model, same context, through `aider`, +fixed a real bug and committed it in 17 seconds. +Swapping in a 12B model produced no answer at all in ten minutes, +because processing that much prompt at 10 tokens per second is simply slow. + +Two practical rules follow. +Shrink the tool surface a local model is shown --- +`--strict-mcp-config --mcp-config '{"mcpServers":{}}'` loads no MCP servers, +and `--allowed-tools` narrows the built-ins. +And tell the harness the real context size, +since Claude Code assumes a 200k window for a model it does not recognize +and would let the conversation grow far past what the model can hold: + +```bash +export CLAUDE_CODE_MAX_CONTEXT_TOKENS=32768 +``` + +For autonomous work on a small local model, +prefer a light harness such as `aider`. +Reserve this route for using a familiar interface offline, +not for getting the best out of the hardware. +::: Air-gapped work aside, the common case is a laptop that is usually online but sometimes is not---on a @@ -376,3 +585,60 @@ This is important when working with: Even with local models, avoid including raw sensitive data in prompts. Work with anonymized or synthetic data wherever possible. + +::: {.callout-important} +#### "Local" is not automatic---some model tags route to a cloud + +Running Ollama does not by itself guarantee that a prompt stays on your machine. +Ollama can serve **cloud-hosted** models alongside local ones, +and those are the models too large to run on a laptop at all, +which is exactly when a tag is tempting. +A cloud-routed tag looks much like a local one in everyday use. + +Two habits keep this honest, +and they matter most in precisely the settings that motivated running locally: + +- **Pull and reference explicitly local tags**, + and treat a tag with no listed download size as cloud-routed until you check + its own page in the [Ollama model library](https://ollama.com/library). +- **Disable the cloud path outright** when the data is regulated, + so the guarantee does not depend on remembering which tag is which: + + ```bash + OLLAMA_NO_CLOUD=1 ollama serve + ``` + +Verify rather than trust either one. +Cut the machine off the network, +or block outbound traffic, +and confirm the agent still completes a real task --- +the check described under +[Verifying you are genuinely offline](#verifying-you-are-genuinely-offline) below. +A setup that quietly depended on a cloud endpoint fails that test immediately. +::: + +#### Verifying you are genuinely offline {#verifying-you-are-genuinely-offline} + +A local setup that has never been tested without a network +is a local setup you are guessing about. +Cutting the machine off entirely is the honest test. +A lighter one that does not disturb the rest of your session +is to make outbound traffic fail for a single command, +while leaving `localhost` reachable: + +```bash +export HTTPS_PROXY=http://127.0.0.1:9 HTTP_PROXY=http://127.0.0.1:9 +export NO_PROXY=localhost,127.0.0.1 + +# Confirm the block is real before trusting the result: +curl -s -m 5 -o /dev/null -w '%{http_code}\n' https://api.anthropic.com # 000 +curl -s -m 5 -o /dev/null -w '%{http_code}\n' http://localhost:11434/api/version # 200 + +aider --yes --message "Fix the off-by-one error in mean()." stats.py +``` + +Check the block itself first, +as above. +A test that passes because the proxy was never applied +tells you nothing, +and looks exactly like success. diff --git a/chapters/ai-tools/small-local-models-agentic.qmd b/chapters/ai-tools/small-local-models-agentic.qmd index c9b1a3f..7932fef 100644 --- a/chapters/ai-tools/small-local-models-agentic.qmd +++ b/chapters/ai-tools/small-local-models-agentic.qmd @@ -60,7 +60,9 @@ enough to run an autonomous loop against: | Family | Sizes worth running locally | License | Best for | |---|---|---|---| -| [Qwen2.5-Coder / Qwen3-Coder](https://ollama.com/library/qwen3-coder) | 7B, 14B, 32B dense; 30B-A3B mixture-of-experts | Apache 2.0 | General-purpose agentic coding across languages | +| [Qwen3-Coder](https://ollama.com/library/qwen3-coder) | 30B-A3B mixture-of-experts (smallest tag, 19 GB) | Apache 2.0 | General-purpose agentic coding across languages | +| [Qwen2.5-Coder](https://ollama.com/library/qwen2.5-coder) | 7B, 14B, 32B dense | Apache 2.0 | Completion and supervised chat. Verify tool calling before trusting it in a loop---see below | +| [Granite 4](https://ollama.com/library/granite4) | 7B-A1B mixture-of-experts (4.2 GB), 32B-A9B | Apache 2.0 | A small, fast tool-caller that fits where the tiers above do not | | [Devstral Small](https://ollama.com/library/devstral) | 24B | Apache 2.0 | Purpose-built for coding agents (multi-file edits, tool use) | | [Codestral](https://ollama.com/library/codestral) | 22B | [Mistral AI Non-Production License](https://mistral.ai/licenses/MNPL-0.1.md) | Fill-in-the-middle completion, not redistribution in a product | | [DeepSeek-Coder-V2](https://ollama.com/library/deepseek-coder-v2) | 16B (Lite) mixture-of-experts | [DeepSeek Model License](https://github.com/deepseek-ai/DeepSeek-Coder-V2/blob/main/LICENSE-MODEL) (commercial use permitted, own terms) | A capable, low-VRAM mixture-of-experts option | @@ -93,6 +95,45 @@ a model is still useful as an assistant you supervise turn by turn (@sec-ai-offline covers exactly that setup), but it is not yet a safe choice to leave unattended. +::: {.callout-important} +#### Parameter count is the wrong first question + +Treat that floor as a statement about *sustained* multi-step loops, +not as a filter to apply before anything else. +Size predicts tool-calling ability poorly enough that checking it first +will mislead you. + +Measured on a 24 GB M2 against a single-function tool schema, +`qwen2.5-coder:14b` produced no tool call at all on four attempts out of four, +returning `finish_reason: stop` each time. +The 4.2 GB `granite4:7b-a1b-h` produced a well-formed call three times out of +three and completed a multi-turn round trip using the result. +A model a third of the size was usable as an agent where the larger one was not, +and no amount of context or prompting fixes a model that will not call a tool. + +The advertised capability list does not settle it either, +and is wrong in both directions: +on the same machine `ollama show qwen2.5-coder:14b` lists `tools` +while the model does not call them, +and `ollama show gemma4:12b` lists none while the model calls them correctly. + +So order the questions this way: + +1. **Does it emit well-formed tool calls?** + One request answers this. + A model that fails here cannot be an agent at any size. +2. **Does it fit, with context?** + Weights plus key-value cache, verified with `ollama ps`. +3. **Is it big enough to sustain a long loop?** + This is where the 24--32B floor applies. + +A small model that clears the first two is worth measuring on your own tasks +before concluding it cannot be left unattended, +because the guardrails below, +not the parameter count, +are what actually bound the damage from a bad step. +::: + #### Hardware tiers | VRAM (or unified memory) | Model tier | Autonomy | @@ -113,6 +154,35 @@ Check the current requirements on the model's own listing tag) rather than a rule of thumb, since quantization schemes change. +::: {.callout-warning} +#### Apple unified memory does not map onto the VRAM column + +Read the unified-memory row as its own scale rather than as the VRAM +figures with a speed penalty attached. +Two deductions come off the headline number before any model loads: + +- **Only about 75% of unified memory is addressable by the GPU** by + default, so a 24 GB machine has roughly 18 GB to work with, not 24. +- **Context is charged on top of the weights.** + Measured on a 24 GB M2, raising a 14B model from 4k to 32k context + moved it from 9.5 GB resident to 15 GB. + +Together those rule out the 30--32B tier on a 24 GB Mac, +even though the headline number matches the VRAM column. +The smallest `qwen3-coder` tag is 19 GB, +which exceeds the addressable ceiling on its own, +leaving nothing for context. +There is no smaller variant of it to fall back to. + +The practical ceiling on 24 GB of unified memory is a **12--14B dense +model at 32k context**, or a mixture-of-experts model of similar +footprint. +Confirm with `ollama ps` after loading: +`PROCESSOR` reading `100% GPU` means it fits, +and anything less means part of the model is on the CPU and the loop +will be far slower than the tier table suggests. +::: + #### Routed architectures: a planner and an executor @sec-ai-offline already shows the mechanics of splitting a task between @@ -223,6 +293,23 @@ own mistake: Keep every prompt, tool call, and gate result the loop produced, so a run that stopped (or that a human later distrusts) can be reviewed after the fact rather than re-run blind. +6. **A tool surface small enough for the model.** + Every tool the harness exposes costs context before the task starts, + because its schema is sent with the prompt, + and a configured Model Context Protocol (MCP) server can add + thousands of tokens of definitions on its own. + Against a 32k local context that overhead competes directly with the + work, + and a small model loses: + asked to read one file, + a 7B model with a full MCP surface loaded produced a paragraph about + missing credentials instead, + and with MCP disabled invented a task list copied from the description + of a tool it had been shown. + Expose the smallest set of tools the step actually needs, + and prefer a light harness over a heavyweight one --- + the same model that failed those attempts fixed a real bug and + committed it in 17 seconds through `aider`. These are the guardrails as a reader-facing rationale. The concrete gate wiring --- diff --git a/inst/WORDLIST b/inst/WORDLIST index 2e240fc..3d07dbd 100644 --- a/inst/WORDLIST +++ b/inst/WORDLIST @@ -63,7 +63,9 @@ Docker codebase Inflexa MCP +Modelfile Ollama +Ollama's OpenAI TORQCLAW TypeScript @@ -75,6 +77,7 @@ proteomics repurposing sandboxed sandboxing +schemas serology transcriptomics translational From 7aee25260185cae3f41061855252ab6f027a44ad Mon Sep 17 00:00:00 2001 From: Douglas Ezra Morrison Date: Sun, 9 Aug 2026 16:26:02 -0700 Subject: [PATCH 2/3] fix: use example.com in the offline probe so link-checker passes api.anthropic.com returns 404 to a plain GET, so lychee rejected it. The probe only needs any external host to demonstrate the block, and a vendor-neutral one states the point better: what matters for an offline claim is that no outbound traffic succeeds. Verified all six external URLs this branch adds now return 200. --- chapters/ai-tools/running-agents-offline.qmd | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/chapters/ai-tools/running-agents-offline.qmd b/chapters/ai-tools/running-agents-offline.qmd index 10fef0a..22c8e00 100644 --- a/chapters/ai-tools/running-agents-offline.qmd +++ b/chapters/ai-tools/running-agents-offline.qmd @@ -631,8 +631,8 @@ export HTTPS_PROXY=http://127.0.0.1:9 HTTP_PROXY=http://127.0.0.1:9 export NO_PROXY=localhost,127.0.0.1 # Confirm the block is real before trusting the result: -curl -s -m 5 -o /dev/null -w '%{http_code}\n' https://api.anthropic.com # 000 -curl -s -m 5 -o /dev/null -w '%{http_code}\n' http://localhost:11434/api/version # 200 +curl -s -m 5 -o /dev/null -w '%{http_code}\n' https://example.com # 000 = blocked +curl -s -m 5 -o /dev/null -w '%{http_code}\n' http://localhost:11434/api/version # 200 = local aider --yes --message "Fix the off-by-one error in mean()." stats.py ``` From 422e319347c6f53dcafb4522584da5ed8c2735d1 Mon Sep 17 00:00:00 2001 From: Douglas Ezra Morrison Date: Sun, 9 Aug 2026 16:43:29 -0700 Subject: [PATCH 3/3] fix: address review findings on PR #57 All four confirmed, none rebutted. 1. Restore the "Falling Back Between Cloud and Local Automatically" heading, which my edit dropped, orphaning the LiteLLM content under the new Claude Code section. Found by reading the diff's deleted lines, which I should have done before the first push. 2. Drop "a third of the size". It was true only of resident size at 32k (4.7 vs 15 GB) and false by parameter count, with no measure stated. Now says "half the parameter count", which is unambiguous. 3. Correct the tool-calling claim, which was wrong. Re-checking with the full `ollama show` output shows gemma4:12b DOES advertise `tools`; my original reading came from `grep -A3`, which truncated the capability list one line before `tools`. The "wrong in both directions" claim is withdrawn. Re-measuring also gave a sharper result than the one it replaces. qwen2.5-coder does not ignore the tools: it emits a correct tool call as prose in `content` while leaving `tool_calls` empty, so a harness sees a plain reply and never runs the tool. The rule is that an advertised capability is necessary but not sufficient. 4. Add HTTP-status checking to the tool-calling snippet. As written it could not distinguish a model declining to call a tool from a request that errored, which is the ambiguity the same file warns about elsewhere. The no-call branch now prints `content`, which is what makes the qwen failure mode visible. Verified against all three branches: granite4 (agent), qwen (no call), bad tag (HTTP 404). --- chapters/ai-tools/running-agents-offline.qmd | 71 ++++++++++++++----- .../ai-tools/small-local-models-agentic.qmd | 27 +++---- inst/WORDLIST | 2 + 3 files changed, 71 insertions(+), 29 deletions(-) diff --git a/chapters/ai-tools/running-agents-offline.qmd b/chapters/ai-tools/running-agents-offline.qmd index 22c8e00..e24b4f8 100644 --- a/chapters/ai-tools/running-agents-offline.qmd +++ b/chapters/ai-tools/running-agents-offline.qmd @@ -79,33 +79,68 @@ Inline completion and chat need only the first. Anything autonomous needs the second, because an agent that cannot call a tool cannot read your repository at all. -The two come apart in practice. +The two come apart in practice, +and the failure is subtler than a model simply refusing. Tested against a single-function tool schema on a 24 GB M2, -`qwen2.5-coder:14b` returned `finish_reason: stop` and no tool call -on four attempts out of four, -while `granite4:7b-a1b-h` returned a well-formed call three times out of three -and completed a full multi-turn round trip --- -despite being a third of the size. - -Do not settle this from the model's advertised capabilities, -which are wrong in both directions. -On the same machine `ollama show qwen2.5-coder:14b` lists `tools` -among its capabilities and the model does not call them, -while `ollama show gemma4:12b` lists no tools and the model calls them correctly. +`qwen2.5-coder:14b` returned `finish_reason: stop` with an empty `tool_calls` +field on four attempts out of four. +It had not ignored the request: +it wrote a correct tool call as ordinary prose in the `content` field, + +```json +{"name": "read_file", "arguments": {"path": "src/main.py"}} +``` + +which is the right JSON in the wrong place. +A harness looks for `tool_calls`, +finds nothing, +and treats the turn as a plain reply, +so the tool never runs. +`granite4:7b-a1b-h`, +at half the parameter count, +returned a well-formed call in the `tool_calls` field three times out of three +and completed a full multi-turn round trip. + +An advertised `tools` capability is necessary but not sufficient, +so do not settle the question with `ollama show`. +On the same machine `qwen2.5-coder:14b` lists `tools` among its capabilities +and still cannot be driven by a harness, +for the reason above. Test it yourself with one request before building a loop on it: ```bash -curl -s http://localhost:11434/v1/chat/completions -H 'content-type: application/json' -d '{ +RESP=$(curl -s -w '\n%{http_code}' \ + http://localhost:11434/v1/chat/completions -H 'content-type: application/json' -d '{ "model": "granite4:7b-a1b-h", "stream": false, "messages": [{"role": "user", "content": "What is in src/main.py? Use the tool."}], "tools": [{"type": "function", "function": {"name": "read_file", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, - "required": ["path"]}}}]}' | grep -o '"tool_calls".*' | head -c 200 + "required": ["path"]}}}]}') + +CODE=$(printf '%s' "$RESP" | tail -1) +BODY=$(printf '%s' "$RESP" | sed '$d') + +if [ "$CODE" != "200" ]; then + echo "request failed (HTTP $CODE): $BODY" # bad tag, tools unsupported, server down +elif printf '%s' "$BODY" | grep -q '"tool_calls"'; then + echo "usable as an agent" +else + echo "no tool call; completion model only" + printf '%s' "$BODY" | grep -o '"content":"[^"]*"' | head -c 200 +fi ``` -Output containing `tool_calls` means the model is usable as an agent. -Silence means it is a completion model, -whatever its model card says. +Check the status code separately from the result. +A request that simply errored --- +a mistyped tag, +a model the server rejects for tool use, +a server that is not running --- +produces the same silence as a model that declined to call the tool, +and only one of those is a fact about the model. +Printing the `content` field on the no-call branch is what distinguishes +a model that ignored the tools +from one that described the call in prose instead of emitting it, +which is the `qwen2.5-coder` case above. ::: **Start the Ollama server:** @@ -385,6 +420,8 @@ Reserve this route for using a familiar interface offline, not for getting the best out of the hardware. ::: +#### Falling Back Between Cloud and Local Automatically + Air-gapped work aside, the common case is a laptop that is usually online but sometimes is not---on a plane, behind a flaky hospital network, or temporarily rate-limited by a cloud provider. diff --git a/chapters/ai-tools/small-local-models-agentic.qmd b/chapters/ai-tools/small-local-models-agentic.qmd index 7932fef..33d736a 100644 --- a/chapters/ai-tools/small-local-models-agentic.qmd +++ b/chapters/ai-tools/small-local-models-agentic.qmd @@ -104,18 +104,21 @@ Size predicts tool-calling ability poorly enough that checking it first will mislead you. Measured on a 24 GB M2 against a single-function tool schema, -`qwen2.5-coder:14b` produced no tool call at all on four attempts out of four, -returning `finish_reason: stop` each time. -The 4.2 GB `granite4:7b-a1b-h` produced a well-formed call three times out of -three and completed a multi-turn round trip using the result. -A model a third of the size was usable as an agent where the larger one was not, -and no amount of context or prompting fixes a model that will not call a tool. - -The advertised capability list does not settle it either, -and is wrong in both directions: -on the same machine `ollama show qwen2.5-coder:14b` lists `tools` -while the model does not call them, -and `ollama show gemma4:12b` lists none while the model calls them correctly. +`qwen2.5-coder:14b` returned an empty `tool_calls` field on four attempts out +of four, with `finish_reason: stop` each time. +It wrote a correct tool call as prose in the `content` field instead, +which no harness will act on. +The 4.2 GB `granite4:7b-a1b-h`, at half the parameter count, +put a well-formed call in `tool_calls` three times out of three and completed +a multi-turn round trip using the result. +The smaller model was usable as an agent where the larger one was not, +and no amount of context or prompting fixes a model whose calls never reach +the field a harness reads. + +The advertised capability list does not settle it either. +`ollama show qwen2.5-coder:14b` lists `tools`, +and the model still cannot be driven by a harness, +so treat the tag as necessary rather than sufficient. So order the questions this way: diff --git a/inst/WORDLIST b/inst/WORDLIST index 3d07dbd..2791c23 100644 --- a/inst/WORDLIST +++ b/inst/WORDLIST @@ -36,12 +36,14 @@ callout ebsite edu emplate +errored gh github glitchy io lintr mergeable +mistyped navbar orking plugin