Local Model Bench 1.10.0 - #8
Merged
Merged
Conversation
Carries two releases. 1.9.0 was prepared and installed locally but never pushed: native MTP for models whose prediction heads ship as a separate file, a sweep across prediction depths, an optional preflight that finds out early when a head will not load, readable `lms` failures, and a text-only twin for benchmarking a vision model without its projector. 1.10.0 is about measuring the machine rather than the measurement. Prompt processing was reported at roughly a third of its real speed. The figure came from LM Studio's prompt_processing interval, which lands within a millisecond or two of time-to-first-token, so it carried the whole fixed cost of a request on top of the prefill; at the prompt sizes the performance tests use, that fixed cost was the measurement. No single request can separate a fixed cost from a per-token one, so each loaded model is now measured over a short prompt and a long one and the difference gives the real per-token rate, naming the overhead that was being charged to the GPU. The older per-request figure is kept beside it. A run can name which llama.cpp build to use, and puts the previous selection back when it ends. This turned out to matter: on an RX 9070 with Gemma 4 12B Q4_K_XL, Vulkan 2.40.0 reached 1027 tok/s of prompt processing against ROCm 2.40.0's 148, and ROCm carried about 16 s of fixed cost before its first token at every prompt size tested. Where two engines have measured the same model, the results tables and graphs will set them side by side. The key/value cache can be quantized per run. It is the part of a run's memory that grows with concurrency rather than with the weights, and on a card whose weights already nearly fill VRAM it decides whether the run measures the GPU or a spill into system RAM. Measured cost: q8_0 for K and V took about 9% of prompt processing. LM Studio exposes no flag for it, so the setting is written to its own per-model configuration for one load and the file restored byte for byte after. Results are also readable one question at a time — all measurements, prompt processing, or token generation — each with the columns that bear on it, and a load that runs out of room now names the settings that decided how much it needed. 233 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| const field=(key:string,q:CacheQuant)=>({key,value:q==='off'?{checked:false,value:'f16'}:{checked:true,value:q}}); | ||
| config.load.fields=[...keep,field(K_CACHE,k),field(V_CACHE,v)]; | ||
| mkdirSync(path.dirname(file),{recursive:true}); | ||
| writeFileSync(file,JSON.stringify(config,null,2)); |
| const keep=(config.load.fields as {key?:unknown}[]).filter(f=>f&&typeof f.key==='string'&&f.key!==SIDECAR&&f.key!==DRAFT_MODEL&&f.key!==MAX_TOKENS&&!EXCLUSIVE.includes(f.key)); | ||
| config.load.fields=[...keep,...EXCLUSIVE.map(key=>({key,value:false})),{key:SIDECAR,value:true},{key:DRAFT_MODEL,value:mtp.draftResource},{key:MAX_TOKENS,value:tokens}]; | ||
| mkdirSync(path.dirname(file),{recursive:true}); | ||
| writeFileSync(file,JSON.stringify(config,null,2)); |
| const raw=(config as any)[key]; | ||
| if(raw===undefined||raw===null)return 'unreported'; | ||
| if(typeof raw==='string')return raw; | ||
| if(typeof raw==='object'&&'checked' in raw)return raw.checked?String(raw.value):'off'; |
Comment on lines
+79
to
+81
| writeFileSync(path.join(plan.folder,MARKER),JSON.stringify({createdBy:'local-model-bench',source:plan.sourceDir,sourceKey:plan.key,created:new Date().toISOString(), | ||
| note:'Every .gguf here is a hard link to the folder named in "source" and uses no extra disk space. Deleting this folder removes the links only; the model itself is untouched.', | ||
| files:plan.files.map(f=>slash(path.relative(plan.folder,f.to)))},null,2)); |
| @@ -0,0 +1,111 @@ | |||
| import type {ChartRow} from './charts'; | |||
| import type {HistoryRow} from './history'; | |||
| import {shortDate,unknownFacet} from './history'; | |||
Quantizing the context made every load fail. llama.cpp cannot use a quantized key/value cache without flash attention, and LM Studio refuses the load outright rather than falling back: Error: V Cache Quantization requires flash attention to be enabled. Please enable flash attention or set V Cache Quantization Type to f16. 1.10.0 wrote the cache fields and assumed flash attention was already on. That assumption does not hold: the default varies by engine and by whatever the model was last loaded with, and on this machine every run that asked for a quantized cache failed within two seconds of pressing Start, which reads as the button doing nothing. A run that asks for a quantized cache now writes the flag that makes it possible, into the same per-model configuration, for the same single load, restored with the rest. A run that leaves the cache alone does not touch the setting. A load that comes back reporting no flash attention is refused before anything is measured, rather than reporting a cache that was not in use. Verified against LM Studio on both engines: the exact case that failed, ROCm 2.40.0 with q4_0 for K and V, now loads with flashAttention, kCacheQuantizationType and vCacheQuantizationType written together and removed together. 236 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The graphs can read across MTP depth. They always plotted against concurrent requests. That is the right axis for a concurrency sweep and the wrong one for a prediction-depth sweep: a sweep run at a single concurrency level gives every depth the same horizontal position, so six depths drew six dots stacked on one vertical line. The depths were there — as six entries in the legend — but nothing separated them on the page, and the run's whole subject was invisible in the charts meant to show it. "Read across" on the Automatic graphs now switches the horizontal axis to maximum predictions, so each model becomes one curve over its depths. The control appears only when a run measured more than one depth. While depth is the axis it stops being folded into the series name, because doing both would split a model into one flat single-point line per depth instead of a curve across them. The tooltip names the depth and still states the concurrency, since a depth curve only means anything at a known concurrency. Nothing changes for a run that did not sweep. Verified against a real completed sweep: on the concurrency axis six rows share one x; on the depth axis they occupy six, joined by a line. 240 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The depth axis existed but nothing opened on it. 1.11.0 added "Read across" and then defaulted it to concurrent requests. For the runs that most need the depth axis — a prediction-depth sweep at a single concurrency level — that meant the graphs still opened with every depth stacked at one horizontal position. The control was there, the feature was not: unless you went looking for a dropdown you had no reason to expect, nothing had changed. The graphs now open on whichever dimension the run actually varied. A run that swept depths without varying concurrency opens on maximum predictions; anything that varied concurrency opens on concurrency, as before. Choosing an axis by hand still overrides it, and the choice is held as null until made so the default follows the run rather than being frozen at first render. The rule is extracted as defaultXAxis() rather than left inline, so it is tested rather than asserted. Checked against the saved runs: the completed six-depth sweep opens on depth with six distinct positions both for all models and narrowed to one; a run whose sweep only reached depth 0 correctly has no depth axis to open on; and the two runs that varied concurrency are untouched. 241 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Off / text-only can now actually load without the projector. The toggle only ever controlled whether an image was sent. The projector loaded regardless, because LM Studio attaches one from the model's index entry and from nowhere else. Rechecked today against LM Studio 2.40.0: none of the 24 llm.load.* keys in its bundle mentions vision, projector or mmproj, and llama-server.exe, llama-server-impl.dll and mtmd.dll carry no --mmproj flag at all, because the engine is driven through LM Studio's own bindings rather than a command line. There is no setting to add. So the toggle now drives the one thing that does work: the text-only copy this app already builds, a tree of hard links holding the same weights and MTP head with no projector, which LM Studio indexes as its own key and reports as non-vision. With Off / text-only selected, any chosen model that would still load a projector is named, with one button to select the copies that exist and another to make the ones that do not. A model with no projector is left alone; a copy LM Studio has not indexed, or whose key it no longer lists, counts as missing rather than being selected and failing to load; and a run that cannot be made fully text-only says so rather than implying otherwise. Also fixes calibrated prompt processing being blank on every sweep. The engine wrote each calibration under one key shape and the results screen read another, so the column never resolved on exactly the runs it was built for. Both sides now use calibrationKey(), which also drops a doubled prefix that spelled depth 0 as "MTP MTP off". Calibrations saved under the old shape are still read. 247 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| // The undepthed key last: a run with no sweep stores exactly one calibration under it. | ||
| for(const key of [...candidates,`prefill:${modelKey}`]){ | ||
| const value=info[key]; | ||
| if(value&&typeof value==='object'&&'marginalTps' in (value as object))return value as PrefillCalibration; |
The three starter JSON tests now accept a fenced answer. subnet, facts and json-extract ask for "only JSON" and say nothing about code fences. A model that returns correct JSON wrapped in ```json was being scored zero for a markdown habit the prompt never mentioned. Measured on Gemma 4 26B A4B: all five subnet values right, scored 0; all four extraction values right, scored 20. Those three set allowCodeFence, and stored copies are migrated once on startup so existing installs pick the change up instead of only new ones. A test edited by hand keeps whatever rules it has. The accounting pack is deliberately excluded. Its prompt says "no prose or code fences", which makes the fence part of what that pack tests, and its assertions that a fenced answer scores zero still hold. So does any test that has not opted in, and a fence that never closes, several fenced blocks, or prose around the JSON are all still failures — a fence is markup around the answer, not prose instead of it. This changes what those three tests score, so a quality comparison spanning the change compares two scoring rules as well as two models. Runs already saved keep the numbers they were given. The core test that asserted the old behaviour is updated rather than removed, and now also pins the thing that must not change: a fenced *wrong* answer earns only the valid-JSON check and nothing for the values inside the fence. 253 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The server does not have to be this machine. The endpoint was restricted to loopback, so LM Studio on another PC, a box on the LAN, or anything else speaking the same API could not be benchmarked from here. Any http:// or https:// address is now accepted. Still refused is anything that is not an origin: credentials in the URL, which belong in the API token field where they are encrypted at rest and never reach the window, and a path, query or fragment, which would change what the address means. Pointing somewhere else changes what this app can honestly claim, so it says so rather than leaving the old promise on screen. The sidebar reads "Remote endpoint" with the host in place of "Your prompts stay on this PC", and the settings field warns that prompts, responses and any token will leave the machine, noting when a plain http address means they travel unencrypted. The badge follows the saved setting rather than the draft, so it describes where runs actually go. The rules move to src/endpoint.ts: the window needs them, and the LM Studio client reaches for node:fs and node:child_process, which cannot be bundled for the renderer. The core test asserting the old loopback-only boundary is updated rather than removed, and still pins what did not change: credentials and paths are refused. 256 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Flash attention is on by default, and a run states it rather than inheriting it. It was only ever turned on as a side effect of quantizing the cache, so every other run took whatever LM Studio happened to have set. That is not a neutral choice. Measured on Gemma 4 26B A4B at f16 cache, changing nothing else: ROCm flash off 260 tok/s prompt processing 7.1 s fixed cost per request ROCm flash on 1281 tok/s 0.06 s Vulkan flash off 277 tok/s 9.2 s Vulkan flash on 1822 tok/s 0.10 s About five times the prompt processing, and the difference between seconds and milliseconds before a first token, on both engines. It also corrects an earlier reading of this project's own runs: ROCm looked slower than Vulkan only because every fast ROCm run happened to quantize its cache, which forced flash attention on. With it on, ROCm generates faster than Vulkan on this card. "Use flash attention" on the Run screen is on unless turned off. It is written to LM Studio's per-model configuration for each load and restored afterwards, like the cache type and the MTP head, and checked against what LM Studio reports before anything is measured: a load that ignored the request is refused rather than measured under a setting nobody chose. Turning it off is supported, because measuring the cost is the only way to know it, and the run says what that will do. A quantized cache with flash attention off is refused up front. llama.cpp cannot do it, and LM Studio's refusal names neither setting. Runs saved before this leave the field unset and were measured under whatever LM Studio had at the time, so they are not evidence either way. Test mocks gain a no-op cache writer, since the config is now written on every load and the adapter exists so tests never touch LM Studio's own files. The test asserting that a non-quantizing run left flash attention alone is rewritten rather than deleted: it now pins that the run's choice wins for the load and the user's setting is restored after. 259 offline tests, typecheck and production build all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Carries two releases. 1.9.0 was prepared and installed locally but never pushed — native MTP for models whose prediction heads ship as a separate file, a sweep across prediction depths, an optional preflight for heads that will not load, readable
lmsfailures, and a text-only twin for benchmarking a vision model without its projector.1.10.0 is about measuring the machine rather than the measurement.
Prompt processing was reported at roughly a third of its real speed
The figure came from LM Studio's
prompt_processinginterval, which in practice lands within a millisecond or two of time-to-first-token — so it carried the whole fixed cost of a request on top of the prefill. At the prompt sizes the performance tests use, that fixed cost was the measurement.No single request can separate a fixed cost from a per-token one, so each loaded model is now measured over a short prompt and a long one; the fixed cost is identical in both and cancels. The older per-request figure is kept beside it rather than silently replaced.
A run can name its llama.cpp build
LM Studio keeps one selected engine and has no per-load switch, so which build produced a set of numbers was not recorded anywhere. A run now selects one and puts the previous choice back when it ends.
This turned out to matter. Measured on an RX 9070 with Gemma 4 12B Q4_K_XL, three prompt sizes, unique prefixes to defeat the prompt cache:
Where two engines have measured the same model, the results tables offer Show ROCm / Show Vulkan / Show all engines and the graphs gain a Compare with menu. Nothing is recalculated: every overlaid point is a saved row, newest run per measurement.
The context can be quantized
The KV cache is the part of a run's memory that grows with concurrency rather than with the weights. K and V are set independently. LM Studio exposes no flag, so the setting is written to its own per-model configuration for one load and the file restored byte for byte afterwards — the same discipline the MTP head uses. Measured cost: q8_0 for both took about 9% of prompt processing.
Also
Verification
233 offline tests, typecheck and production build all pass. The engine comparison, calibration and cache quantization were exercised end to end against both backends on real hardware, and the resulting rows checked through the same code the UI uses.
Not verified: the rendered UI has not been clicked through: the logic layer was checked against the real 217-row history and a purpose-built two-engine dataset.
🤖 Generated with Claude Code