Skip to content

Local Model Bench 1.10.0 - #8

Merged
M007-Net merged 8 commits into
mainfrom
release/1.10.0
Sep 18, 2026
Merged

M007-Net merged 8 commits into
mainfrom
release/1.10.0

Conversation

@M007-Net

Copy link
Copy Markdown
Owner

Carries two releases. 1.9.0 was prepared and installed locally but never pushed — native MTP for models whose prediction heads ship as a separate file, a sweep across prediction depths, an optional preflight for heads that will not load, readable lms failures, and a text-only twin for benchmarking a vision model without its projector.

1.10.0 is about measuring the machine rather than the measurement.

Prompt processing was reported at roughly a third of its real speed

The figure came from LM Studio's prompt_processing interval, which in practice lands within a millisecond or two of time-to-first-token — so it carried the whole fixed cost of a request on top of the prefill. At the prompt sizes the performance tests use, that fixed cost was the measurement.

No single request can separate a fixed cost from a per-token one, so each loaded model is now measured over a short prompt and a long one; the fixed cost is identical in both and cancels. The older per-request figure is kept beside it rather than silently replaced.

A run can name its llama.cpp build

LM Studio keeps one selected engine and has no per-load switch, so which build produced a set of numbers was not recorded anywhere. A run now selects one and puts the previous choice back when it ends.

This turned out to matter. Measured on an RX 9070 with Gemma 4 12B Q4_K_XL, three prompt sizes, unique prefixes to defeat the prompt cache:

Vulkan 2.40.0 ROCm 2.40.0
Prompt processing 1027 tok/s 148 tok/s
Fixed cost before first token none measurable ~16 s
Generation 61 tok/s 51 tok/s

Where two engines have measured the same model, the results tables offer Show ROCm / Show Vulkan / Show all engines and the graphs gain a Compare with menu. Nothing is recalculated: every overlaid point is a saved row, newest run per measurement.

The context can be quantized

The KV cache is the part of a run's memory that grows with concurrency rather than with the weights. K and V are set independently. LM Studio exposes no flag, so the setting is written to its own per-model configuration for one load and the file restored byte for byte afterwards — the same discipline the MTP head uses. Measured cost: q8_0 for both took about 9% of prompt processing.

Also

  • Results read one question at a time: all measurements, prompt processing, or token generation, each with the columns that bear on it.
  • A load that runs out of room now names the settings that decided how much it needed, instead of only LM Studio's own wording.

Verification

233 offline tests, typecheck and production build all pass. The engine comparison, calibration and cache quantization were exercised end to end against both backends on real hardware, and the resulting rows checked through the same code the UI uses.

Not verified: the rendered UI has not been clicked through: the logic layer was checked against the real 217-row history and a purpose-built two-engine dataset.

🤖 Generated with Claude Code

Carries two releases. 1.9.0 was prepared and installed locally but never pushed:
native MTP for models whose prediction heads ship as a separate file, a sweep
across prediction depths, an optional preflight that finds out early when a head
will not load, readable `lms` failures, and a text-only twin for benchmarking a
vision model without its projector.

1.10.0 is about measuring the machine rather than the measurement.

Prompt processing was reported at roughly a third of its real speed. The figure
came from LM Studio's prompt_processing interval, which lands within a millisecond
or two of time-to-first-token, so it carried the whole fixed cost of a request on
top of the prefill; at the prompt sizes the performance tests use, that fixed cost
was the measurement. No single request can separate a fixed cost from a per-token
one, so each loaded model is now measured over a short prompt and a long one and
the difference gives the real per-token rate, naming the overhead that was being
charged to the GPU. The older per-request figure is kept beside it.

A run can name which llama.cpp build to use, and puts the previous selection back
when it ends. This turned out to matter: on an RX 9070 with Gemma 4 12B Q4_K_XL,
Vulkan 2.40.0 reached 1027 tok/s of prompt processing against ROCm 2.40.0's 148,
and ROCm carried about 16 s of fixed cost before its first token at every prompt
size tested. Where two engines have measured the same model, the results tables
and graphs will set them side by side.

The key/value cache can be quantized per run. It is the part of a run's memory
that grows with concurrency rather than with the weights, and on a card whose
weights already nearly fill VRAM it decides whether the run measures the GPU or a
spill into system RAM. Measured cost: q8_0 for K and V took about 9% of prompt
processing. LM Studio exposes no flag for it, so the setting is written to its own
per-model configuration for one load and the file restored byte for byte after.

Results are also readable one question at a time — all measurements, prompt
processing, or token generation — each with the columns that bear on it, and a
load that runs out of room now names the settings that decided how much it needed.

233 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comment thread electron/cache-quant.ts
const field=(key:string,q:CacheQuant)=>({key,value:q==='off'?{checked:false,value:'f16'}:{checked:true,value:q}});
config.load.fields=[...keep,field(K_CACHE,k),field(V_CACHE,v)];
mkdirSync(path.dirname(file),{recursive:true});
writeFileSync(file,JSON.stringify(config,null,2));
Comment thread electron/mtp-sidecar.ts
const keep=(config.load.fields as {key?:unknown}[]).filter(f=>f&&typeof f.key==='string'&&f.key!==SIDECAR&&f.key!==DRAFT_MODEL&&f.key!==MAX_TOKENS&&!EXCLUSIVE.includes(f.key));
config.load.fields=[...keep,...EXCLUSIVE.map(key=>({key,value:false})),{key:SIDECAR,value:true},{key:DRAFT_MODEL,value:mtp.draftResource},{key:MAX_TOKENS,value:tokens}];
mkdirSync(path.dirname(file),{recursive:true});
writeFileSync(file,JSON.stringify(config,null,2));
Comment thread electron/cache-quant.ts
const raw=(config as any)[key];
if(raw===undefined||raw===null)return 'unreported';
if(typeof raw==='string')return raw;
if(typeof raw==='object'&&'checked' in raw)return raw.checked?String(raw.value):'off';
Comment thread electron/text-only.ts
Comment on lines +79 to +81
writeFileSync(path.join(plan.folder,MARKER),JSON.stringify({createdBy:'local-model-bench',source:plan.sourceDir,sourceKey:plan.key,created:new Date().toISOString(),
note:'Every .gguf here is a hard link to the folder named in "source" and uses no extra disk space. Deleting this folder removes the links only; the model itself is untouched.',
files:plan.files.map(f=>slash(path.relative(plan.folder,f.to)))},null,2));
Comment thread src/prefill.ts Fixed
Comment thread src/chart-compare.ts
@@ -0,0 +1,111 @@
import type {ChartRow} from './charts';
import type {HistoryRow} from './history';
import {shortDate,unknownFacet} from './history';
M007-Net and others added 4 commits September 17, 2026 03:25
Quantizing the context made every load fail.

llama.cpp cannot use a quantized key/value cache without flash attention, and
LM Studio refuses the load outright rather than falling back:

  Error: V Cache Quantization requires flash attention to be enabled.
  Please enable flash attention or set V Cache Quantization Type to f16.

1.10.0 wrote the cache fields and assumed flash attention was already on. That
assumption does not hold: the default varies by engine and by whatever the model
was last loaded with, and on this machine every run that asked for a quantized
cache failed within two seconds of pressing Start, which reads as the button
doing nothing.

A run that asks for a quantized cache now writes the flag that makes it possible,
into the same per-model configuration, for the same single load, restored with the
rest. A run that leaves the cache alone does not touch the setting. A load that
comes back reporting no flash attention is refused before anything is measured,
rather than reporting a cache that was not in use.

Verified against LM Studio on both engines: the exact case that failed, ROCm 2.40.0
with q4_0 for K and V, now loads with flashAttention, kCacheQuantizationType and
vCacheQuantizationType written together and removed together.

236 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The graphs can read across MTP depth.

They always plotted against concurrent requests. That is the right axis for a
concurrency sweep and the wrong one for a prediction-depth sweep: a sweep run at a
single concurrency level gives every depth the same horizontal position, so six
depths drew six dots stacked on one vertical line. The depths were there — as six
entries in the legend — but nothing separated them on the page, and the run's whole
subject was invisible in the charts meant to show it.

"Read across" on the Automatic graphs now switches the horizontal axis to maximum
predictions, so each model becomes one curve over its depths. The control appears
only when a run measured more than one depth. While depth is the axis it stops being
folded into the series name, because doing both would split a model into one flat
single-point line per depth instead of a curve across them. The tooltip names the
depth and still states the concurrency, since a depth curve only means anything at a
known concurrency. Nothing changes for a run that did not sweep.

Verified against a real completed sweep: on the concurrency axis six rows share one
x; on the depth axis they occupy six, joined by a line.

240 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The depth axis existed but nothing opened on it.

1.11.0 added "Read across" and then defaulted it to concurrent requests. For the runs
that most need the depth axis — a prediction-depth sweep at a single concurrency level
— that meant the graphs still opened with every depth stacked at one horizontal
position. The control was there, the feature was not: unless you went looking for a
dropdown you had no reason to expect, nothing had changed.

The graphs now open on whichever dimension the run actually varied. A run that swept
depths without varying concurrency opens on maximum predictions; anything that varied
concurrency opens on concurrency, as before. Choosing an axis by hand still overrides
it, and the choice is held as null until made so the default follows the run rather
than being frozen at first render.

The rule is extracted as defaultXAxis() rather than left inline, so it is tested
rather than asserted. Checked against the saved runs: the completed six-depth sweep
opens on depth with six distinct positions both for all models and narrowed to one;
a run whose sweep only reached depth 0 correctly has no depth axis to open on; and
the two runs that varied concurrency are untouched.

241 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Off / text-only can now actually load without the projector.

The toggle only ever controlled whether an image was sent. The projector loaded
regardless, because LM Studio attaches one from the model's index entry and from
nowhere else. Rechecked today against LM Studio 2.40.0: none of the 24 llm.load.*
keys in its bundle mentions vision, projector or mmproj, and llama-server.exe,
llama-server-impl.dll and mtmd.dll carry no --mmproj flag at all, because the engine
is driven through LM Studio's own bindings rather than a command line. There is no
setting to add.

So the toggle now drives the one thing that does work: the text-only copy this app
already builds, a tree of hard links holding the same weights and MTP head with no
projector, which LM Studio indexes as its own key and reports as non-vision. With
Off / text-only selected, any chosen model that would still load a projector is
named, with one button to select the copies that exist and another to make the ones
that do not. A model with no projector is left alone; a copy LM Studio has not
indexed, or whose key it no longer lists, counts as missing rather than being
selected and failing to load; and a run that cannot be made fully text-only says so
rather than implying otherwise.

Also fixes calibrated prompt processing being blank on every sweep. The engine wrote
each calibration under one key shape and the results screen read another, so the
column never resolved on exactly the runs it was built for. Both sides now use
calibrationKey(), which also drops a doubled prefix that spelled depth 0 as
"MTP MTP off". Calibrations saved under the old shape are still read.

247 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comment thread src/prefill.ts
// The undepthed key last: a run with no sweep stores exactly one calibration under it.
for(const key of [...candidates,`prefill:${modelKey}`]){
const value=info[key];
if(value&&typeof value==='object'&&'marginalTps' in (value as object))return value as PrefillCalibration;
M007-Net and others added 3 commits September 17, 2026 06:59
The three starter JSON tests now accept a fenced answer.

subnet, facts and json-extract ask for "only JSON" and say nothing about code
fences. A model that returns correct JSON wrapped in ```json was being scored zero
for a markdown habit the prompt never mentioned. Measured on Gemma 4 26B A4B: all
five subnet values right, scored 0; all four extraction values right, scored 20.

Those three set allowCodeFence, and stored copies are migrated once on startup so
existing installs pick the change up instead of only new ones. A test edited by hand
keeps whatever rules it has.

The accounting pack is deliberately excluded. Its prompt says "no prose or code
fences", which makes the fence part of what that pack tests, and its assertions that
a fenced answer scores zero still hold. So does any test that has not opted in, and
a fence that never closes, several fenced blocks, or prose around the JSON are all
still failures — a fence is markup around the answer, not prose instead of it.

This changes what those three tests score, so a quality comparison spanning the
change compares two scoring rules as well as two models. Runs already saved keep the
numbers they were given.

The core test that asserted the old behaviour is updated rather than removed, and now
also pins the thing that must not change: a fenced *wrong* answer earns only the
valid-JSON check and nothing for the values inside the fence.

253 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The server does not have to be this machine.

The endpoint was restricted to loopback, so LM Studio on another PC, a box on the LAN,
or anything else speaking the same API could not be benchmarked from here. Any http://
or https:// address is now accepted. Still refused is anything that is not an origin:
credentials in the URL, which belong in the API token field where they are encrypted at
rest and never reach the window, and a path, query or fragment, which would change what
the address means.

Pointing somewhere else changes what this app can honestly claim, so it says so rather
than leaving the old promise on screen. The sidebar reads "Remote endpoint" with the
host in place of "Your prompts stay on this PC", and the settings field warns that
prompts, responses and any token will leave the machine, noting when a plain http
address means they travel unencrypted. The badge follows the saved setting rather than
the draft, so it describes where runs actually go.

The rules move to src/endpoint.ts: the window needs them, and the LM Studio client
reaches for node:fs and node:child_process, which cannot be bundled for the renderer.

The core test asserting the old loopback-only boundary is updated rather than removed,
and still pins what did not change: credentials and paths are refused.

256 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Flash attention is on by default, and a run states it rather than inheriting it.

It was only ever turned on as a side effect of quantizing the cache, so every other run
took whatever LM Studio happened to have set. That is not a neutral choice. Measured on
Gemma 4 26B A4B at f16 cache, changing nothing else:

  ROCm   flash off   260 tok/s prompt processing   7.1 s fixed cost per request
  ROCm   flash on   1281 tok/s                     0.06 s
  Vulkan flash off   277 tok/s                     9.2 s
  Vulkan flash on   1822 tok/s                     0.10 s

About five times the prompt processing, and the difference between seconds and
milliseconds before a first token, on both engines. It also corrects an earlier reading
of this project's own runs: ROCm looked slower than Vulkan only because every fast ROCm
run happened to quantize its cache, which forced flash attention on. With it on, ROCm
generates faster than Vulkan on this card.

"Use flash attention" on the Run screen is on unless turned off. It is written to LM
Studio's per-model configuration for each load and restored afterwards, like the cache
type and the MTP head, and checked against what LM Studio reports before anything is
measured: a load that ignored the request is refused rather than measured under a
setting nobody chose. Turning it off is supported, because measuring the cost is the
only way to know it, and the run says what that will do.

A quantized cache with flash attention off is refused up front. llama.cpp cannot do it,
and LM Studio's refusal names neither setting.

Runs saved before this leave the field unset and were measured under whatever LM Studio
had at the time, so they are not evidence either way.

Test mocks gain a no-op cache writer, since the config is now written on every load and
the adapter exists so tests never touch LM Studio's own files. The test asserting that a
non-quantizing run left flash attention alone is rewritten rather than deleted: it now
pins that the run's choice wins for the load and the user's setting is restored after.

259 offline tests, typecheck and production build all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@M007-Net
M007-Net merged commit 1952a21 into main Sep 18, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants