From b36cdb646589efb90ca1380c4b4e6c78c5cb55e0 Mon Sep 17 00:00:00 2001 From: Agent Date: Fri, 7 Aug 2026 12:27:22 +0200 Subject: [PATCH 1/2] docs(speculative): add measured Pascal sm_61 MTP results and two cautions Adds docs/benchmarks/pascal-p5200-mtp.md: a baseline-vs-`draft-mtp` comparison on a Quadro P5200 (compute capability 6.1) with Gemma 4 12B QAT, plus the 16 raw artifacts the numbers were computed from. Measured: 24.02 -> 56.44 tok/s generation (2.35x) for +300 MiB VRAM and +23 W, at 82.88% draft acceptance (276/333). Prompt processing unchanged. Two Pascal cautions that are easy to hit and are not obvious from the option list: - `-md` without an explicit `--spec-type draft-mtp` does not enable MTP. The server starts and serves normally, so the failure is silent; the tell is `draft_n == 0` in the response timings. - `--spec-draft-backend-sampling` defaults to enabled, and with it enabled on this sm_61 device repeated greedy requests produced different output. `--no-spec-draft-backend-sampling` restored run-to-run repeatability at no material speed cost. Reported as an observation on one device, not a diagnosis; not bisected, no claim about other architectures. Also documents `--spec-draft-backend-sampling` in the draft-model option list in docs/speculative.md, where it was missing, and links the new page from the Benchmarking section. The write-up is explicit that MTP output is not bit-identical to non-speculative output (batched target evaluation can pick a different greedy token than single-token evaluation), and that the 2.35x figure was taken at `--parallel 1` and is therefore an upper bound for a multi-slot server. Docs and data only; no code or build changes. Nothing is modified or removed - the two edits to docs/speculative.md are pure insertions. Acting agent: crystal-mom --- .../data/pascal-p5200-2026-08-02/README.md | 142 +++++++++++ .../pascal-p5200-2026-08-02/gemma-bench.jsonl | 4 + .../pascal-p5200-2026-08-02/gemma-bench.log | 22 ++ .../gemma-telemetry.csv | 77 ++++++ .../gemma12-baseline-bench.jsonl | 3 + .../gemma12-baseline-smoke.json | 1 + .../gemma12-baseline-telemetry.csv | 34 +++ .../gemma12-mtp-bench.jsonl | 3 + .../gemma12-mtp-cpu-sampler-1.json | 1 + .../gemma12-mtp-cpu-sampler-2.json | 1 + .../gemma12-mtp-smoke-2.json | 1 + .../gemma12-mtp-smoke.json | 1 + .../gemma12-mtp-telemetry.csv | 16 ++ .../pascal-p5200-2026-08-02/qwen-bench.jsonl | 4 + .../pascal-p5200-2026-08-02/qwen-bench.log | 22 ++ .../qwen-telemetry.csv | 239 ++++++++++++++++++ docs/benchmarks/pascal-p5200-mtp.md | 123 +++++++++ docs/speculative.md | 5 + 18 files changed, 699 insertions(+) create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/README.md create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.jsonl create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.log create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-telemetry.csv create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-bench.jsonl create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-smoke.json create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-telemetry.csv create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-bench.jsonl create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-1.json create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-2.json create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke-2.json create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke.json create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-telemetry.csv create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.jsonl create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.log create mode 100644 docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-telemetry.csv create mode 100644 docs/benchmarks/pascal-p5200-mtp.md diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/README.md b/docs/benchmarks/data/pascal-p5200-2026-08-02/README.md new file mode 100644 index 000000000000..f9a6dfcb80ec --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/README.md @@ -0,0 +1,142 @@ +# Pascal CUDA benchmark: Qwen3.6-35B-A3B and Gemma 4 + +Tested on 2026-08-02 with the `ht` branch of ht-llama.cpp. + +## Runtime + +- ht-llama.cpp commit: `798cf6cbe56440132df23eb2318f587b16e3c00e` +- build: `b9862-798cf6cbe` +- GPU: NVIDIA Quadro P5200, 16 GiB, compute capability 6.1 +- CPU: Intel Core i7-7820HQ, 4 cores / 8 threads +- CUDA build settings: + - `CMAKE_CUDA_ARCHITECTURES=61` + - `GGML_CUDA_FORCE_MMQ=ON` + - `GGML_CUDA_F16=OFF` +- benchmark settings: 4 CPU threads, flash attention on, automatic device fitting, + 1024 MiB VRAM safety target, 4096-token fit context, three measured repetitions + +## Verified models + +| Model | File size | SHA-256 | +| --- | ---: | --- | +| Qwen3.6-35B-A3B Q4_K_M | 20,419,565,568 bytes | `671e47e0ec53c665d048b98c3ecbfd5236b5ca9c3e02ed19fc8f81f7b85140c7` | +| Gemma 4 26B-A4B IT QAT Q4_0 | 14,439,363,584 bytes | `3eca3b8f6d7baf218a7dd6bba5fb59a56ee25fe2d567b6f5f589b4f697eca51d` | + +Both files passed `sha256sum --check` before testing. + +## Results + +| Model | PP128 | PP512 | TG32 | TG128 | Warm TG128 | +| --- | ---: | ---: | ---: | ---: | ---: | +| Qwen3.6-35B-A3B Q4_K_M | 225.72 tok/s | 445.85 tok/s | 30.35 tok/s | 27.50 tok/s | 33.73 tok/s | +| Gemma 4 26B-A4B Q4_0 | 503.47 tok/s | 861.37 tok/s | 52.23 tok/s | 51.52 tok/s | 51.80 tok/s | + +`PP` is prompt processing and `TG` is token generation. The normal result columns +are arithmetic means of all three samples. Qwen's first TG128 sample was cold-page +limited at 15.06 tok/s; its next two samples were 33.12 and 34.33 tok/s, whose mean +is shown as Warm TG128. Gemma's TG128 samples were stable at 50.95, 51.84, and +51.76 tok/s. + +## Device placement and telemetry + +| Model | Placement | Peak VRAM | Peak GPU | Peak power | Peak temperature | +| --- | --- | ---: | ---: | ---: | ---: | +| Qwen3.6 | 41/41 layers use CUDA; 13 layers partially overflow to host; 14,146 MiB CUDA model buffer | 15,036 MiB | 100% | 161.82 W | 64 C | +| Gemma 4 | 31/31 layers fully offloaded; 13,755 MiB CUDA model buffer; 577.5 MiB CPU-mapped model data | 14,652 MiB | 100% | 155.68 W | 66 C | + +Across samples with more than 1 GiB VRAM allocated, average GPU utilization was +30.9% for Qwen and 83.9% for Gemma. Qwen's host overflow and file paging account +for the lower utilization and larger cold-run variance. Gemma fits almost entirely +in VRAM and is correspondingly steadier. + +## Smoke tests + +Both models loaded, generated tokens, and exited with status 0 using the same +Pascal CUDA build. Single-turn checks produced `QWEN_SMOKE_OK` and +`GEMMA_SMOKE_OK`. Gemma's embedded chat template requires `--jinja`; without it, +the legacy template path reports that the custom template is unsupported. + +Representative command: + +```sh +./build-cuda/bin/llama-bench \ + -m /home/me/Models/Qwen3.6-35B-A3B-Q4_K_M.gguf \ + -p 128,512 -n 32,128 -r 3 -t 4 \ + -fa on -fitt 1024 -fitc 4096 --progress -o jsonl +``` + +For Gemma chat/completion, add `--jinja`: + +```sh +./build-cuda/bin/llama completion \ + -m /home/me/Models/gemma-4-26B_q4_0-it.gguf \ + -c 2048 -ngl auto -fit on -fitt 1024 -fa on --jinja -st \ + -p 'Your prompt' -n 256 +``` + +Raw data: + +- `qwen-bench.jsonl` and `gemma-bench.jsonl`: llama-bench results and samples +- `qwen-telemetry.csv` and `gemma-telemetry.csv`: 500 ms `nvidia-smi` samples +- `qwen-bench.log` and `gemma-bench.log`: benchmark progress and device detection + +## Gemma 4 12B official QAT + MTP + +The official Google QAT target and the matching QAT-derived MTP assistant were +also installed and tested with the same `ht` Pascal build: + +| Role | File | Size | SHA-256 | +| --- | --- | ---: | --- | +| Target | `gemma-4-12b-it-qat-q4_0.gguf` | 6,975,879,296 bytes | `93567e57a8fe10b23569b9d9ec38cd005deedf71e29477c421a4b83f418a538b` | +| MTP assistant | `mtp-gemma-4-12B-it-Q4_0.gguf` | 253,708,960 bytes | `b894e614824dfc2746b26d3c3ba78c50000a464382682502392b4325257b7602` | + +Both files were exact-size checked and SHA-256 verified after download. The target +is from `google/gemma-4-12B-it-qat-q4_0-gguf`; the assistant is the QAT-derived +`gemma4-assistant` GGUF from `ggml-org/gemma-4-12B-it-GGUF`. + +### Server generation benchmark + +The test used one server slot, a 4096-token context, full GPU offload, flash +attention, greedy sampling, and three independent 128-token requests. + +| Mode | Generation | Prompt processing | Peak VRAM | Peak GPU | Peak power | +| --- | ---: | ---: | ---: | ---: | ---: | +| Target only | 24.02 tok/s | 159.77 tok/s | 7,524 MiB | 98% | 149.10 W | +| QAT MTP | 56.44 tok/s | 156.16 tok/s | 7,824 MiB | 94% | 172.38 W | + +MTP improved generation throughput by **2.35x**. Every MTP run accepted 92 of +111 proposed draft tokens (82.88%), and all three MTP responses were byte-identical +to one another. All three target-only responses were also byte-identical to one +another. + +The required MTP launch options are: + +```sh +./build-cuda/bin/llama serve \ + -m /home/me/Models/gemma-4-12b-it-qat-q4_0.gguf \ + -md /home/me/Models/mtp-gemma-4-12B-it-Q4_0.gguf \ + -c 4096 -ngl all -ngld all -fa on --parallel 1 \ + --spec-type draft-mtp --spec-draft-n-max 16 --spec-draft-p-min 0.9 \ + --no-spec-draft-backend-sampling \ + --host 127.0.0.1 --port 8080 --no-webui --jinja +``` + +Two Pascal-specific cautions were confirmed. Supplying `-md` without the explicit +`--spec-type draft-mtp` does not enable MTP. Also, draft backend sampling produced +non-repeatable greedy output on this P5200; `--no-spec-draft-backend-sampling` +made repeated MTP runs deterministic without materially reducing speed. + +The deterministic MTP response was not byte-identical to the target-only response. +The target still verifies proposed tokens, but batched target evaluation can choose +a different greedy token from single-token evaluation because GPU floating-point +evaluation order differs. Therefore this local result demonstrates verified-token +speculation and repeatability, but it should not be described as bit-identical to +non-speculative generation. + +Additional raw data: + +- `gemma12-baseline-bench.jsonl` and `gemma12-mtp-bench.jsonl`: three API responses + per mode, including llama.cpp timing and MTP acceptance counters +- `gemma12-baseline-telemetry.csv` and `gemma12-mtp-telemetry.csv`: 500 ms GPU samples +- `gemma12-*-smoke.json` and `gemma12-mtp-cpu-sampler-*.json`: smoke and + determinism-isolation runs diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.jsonl b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.jsonl new file mode 100644 index 000000000000..7cf9bad44659 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.jsonl @@ -0,0 +1,4 @@ +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/gemma-4-26B_q4_0-it.gguf", "model_type": "gemma4 26B.A4B Q4_0", "model_size": 14423538808, "model_n_params": 25233142046, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 128, "n_gen": 0, "n_depth": 0, "test_time": "2026-08-02T11:00:11Z", "avg_ns": 254244569, "stddev_ns": 1942637, "avg_ts": 503.471781, "stddev_ts": 3.834547, "samples_ns": [ 253634389, 256419044, 252680274 ],"samples_ts": [ 504.663, 499.183, 506.569 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/gemma-4-26B_q4_0-it.gguf", "model_type": "gemma4 26B.A4B Q4_0", "model_size": 14423538808, "model_n_params": 25233142046, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 512, "n_gen": 0, "n_depth": 0, "test_time": "2026-08-02T11:00:17Z", "avg_ns": 594408106, "stddev_ns": 2621377, "avg_ts": 861.372253, "stddev_ts": 3.805618, "samples_ns": [ 596507863, 591470453, 595246004 ],"samples_ts": [ 858.329, 865.639, 860.149 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/gemma-4-26B_q4_0-it.gguf", "model_type": "gemma4 26B.A4B Q4_0", "model_size": 14423538808, "model_n_params": 25233142046, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 0, "n_gen": 32, "n_depth": 0, "test_time": "2026-08-02T11:00:25Z", "avg_ns": 612631473, "stddev_ns": 913033, "avg_ts": 52.233763, "stddev_ts": 0.077763, "samples_ns": [ 611930312, 612300661, 613663447 ],"samples_ts": [ 52.2935, 52.2619, 52.1458 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/gemma-4-26B_q4_0-it.gguf", "model_type": "gemma4 26B.A4B Q4_0", "model_size": 14423538808, "model_n_params": 25233142046, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 0, "n_gen": 128, "n_depth": 0, "test_time": "2026-08-02T11:00:31Z", "avg_ns": 2484774465, "stddev_ns": 23834946, "avg_ts": 51.516873, "stddev_ts": 0.491517, "samples_ns": [ 2512217852, 2469251358, 2472854187 ],"samples_ts": [ 50.951, 51.8376, 51.762 ]} diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.log b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.log new file mode 100644 index 000000000000..e24d4dfa87f3 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-bench.log @@ -0,0 +1,22 @@ +ggml_cuda_init: found 1 CUDA devices (Total VRAM: 16266 MiB): + Device 0: Quadro P5200, compute capability 6.1, VMM: yes, VRAM: 16266 MiB +llama-bench: benchmark 1/4: starting +llama-bench: benchmark 1/4: warmup prompt run +llama-bench: benchmark 1/4: prompt run 1/3 +llama-bench: benchmark 1/4: prompt run 2/3 +llama-bench: benchmark 1/4: prompt run 3/3 +llama-bench: benchmark 2/4: starting +llama-bench: benchmark 2/4: warmup prompt run +llama-bench: benchmark 2/4: prompt run 1/3 +llama-bench: benchmark 2/4: prompt run 2/3 +llama-bench: benchmark 2/4: prompt run 3/3 +llama-bench: benchmark 3/4: starting +llama-bench: benchmark 3/4: warmup generation run +llama-bench: benchmark 3/4: generation run 1/3 +llama-bench: benchmark 3/4: generation run 2/3 +llama-bench: benchmark 3/4: generation run 3/3 +llama-bench: benchmark 4/4: starting +llama-bench: benchmark 4/4: warmup generation run +llama-bench: benchmark 4/4: generation run 1/3 +llama-bench: benchmark 4/4: generation run 2/3 +llama-bench: benchmark 4/4: generation run 3/3 diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-telemetry.csv b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-telemetry.csv new file mode 100644 index 000000000000..a22a10cea435 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma-telemetry.csv @@ -0,0 +1,77 @@ +timestamp, utilization.gpu [%], memory.used [MiB], power.draw [W], temperature.gpu +2026/08/02 13:00:01.733, 0 %, 90 MiB, 16.44 W, 54 +2026/08/02 13:00:02.237, 0 %, 92 MiB, 9.72 W, 54 +2026/08/02 13:00:02.739, 0 %, 198 MiB, 40.06 W, 55 +2026/08/02 13:00:03.241, 0 %, 198 MiB, 40.23 W, 55 +2026/08/02 13:00:03.744, 0 %, 198 MiB, 40.20 W, 55 +2026/08/02 13:00:04.246, 0 %, 198 MiB, 40.19 W, 55 +2026/08/02 13:00:04.748, 0 %, 198 MiB, 40.30 W, 55 +2026/08/02 13:00:05.249, 0 %, 198 MiB, 40.33 W, 55 +2026/08/02 13:00:05.751, 0 %, 198 MiB, 40.31 W, 55 +2026/08/02 13:00:06.252, 0 %, 198 MiB, 40.34 W, 55 +2026/08/02 13:00:06.754, 0 %, 198 MiB, 40.17 W, 55 +2026/08/02 13:00:07.256, 0 %, 198 MiB, 40.30 W, 55 +2026/08/02 13:00:07.757, 0 %, 198 MiB, 40.31 W, 55 +2026/08/02 13:00:08.259, 0 %, 198 MiB, 40.31 W, 55 +2026/08/02 13:00:08.760, 0 %, 198 MiB, 40.19 W, 55 +2026/08/02 13:00:09.262, 32 %, 13954 MiB, 41.76 W, 55 +2026/08/02 13:00:09.764, 84 %, 13954 MiB, 54.60 W, 56 +2026/08/02 13:00:10.265, 86 %, 13954 MiB, 54.60 W, 56 +2026/08/02 13:00:10.767, 86 %, 13954 MiB, 54.76 W, 56 +2026/08/02 13:00:11.269, 1 %, 14156 MiB, 62.33 W, 56 +2026/08/02 13:00:11.771, 100 %, 14178 MiB, 145.38 W, 60 +2026/08/02 13:00:12.273, 100 %, 14178 MiB, 148.64 W, 62 +2026/08/02 13:00:12.774, 63 %, 208 MiB, 51.67 W, 59 +2026/08/02 13:00:13.276, 0 %, 208 MiB, 51.31 W, 58 +2026/08/02 13:00:13.777, 0 %, 208 MiB, 51.18 W, 58 +2026/08/02 13:00:14.279, 0 %, 208 MiB, 51.18 W, 58 +2026/08/02 13:00:14.781, 0 %, 208 MiB, 51.14 W, 58 +2026/08/02 13:00:15.282, 0 %, 208 MiB, 51.14 W, 58 +2026/08/02 13:00:15.784, 84 %, 13964 MiB, 53.61 W, 58 +2026/08/02 13:00:16.285, 62 %, 13964 MiB, 53.30 W, 57 +2026/08/02 13:00:16.787, 56 %, 13964 MiB, 52.89 W, 58 +2026/08/02 13:00:17.289, 68 %, 13964 MiB, 53.34 W, 58 +2026/08/02 13:00:17.790, 0 %, 13964 MiB, 51.04 W, 57 +2026/08/02 13:00:18.292, 100 %, 14652 MiB, 142.02 W, 61 +2026/08/02 13:00:18.793, 100 %, 14652 MiB, 142.68 W, 62 +2026/08/02 13:00:19.295, 97 %, 14652 MiB, 125.08 W, 63 +2026/08/02 13:00:19.797, 97 %, 14652 MiB, 155.68 W, 64 +2026/08/02 13:00:20.314, 100 %, 14652 MiB, 148.33 W, 64 +2026/08/02 13:00:20.815, 67 %, 208 MiB, 52.46 W, 61 +2026/08/02 13:00:21.317, 0 %, 208 MiB, 51.96 W, 60 +2026/08/02 13:00:21.818, 0 %, 208 MiB, 51.67 W, 60 +2026/08/02 13:00:22.320, 0 %, 208 MiB, 51.63 W, 59 +2026/08/02 13:00:22.821, 0 %, 208 MiB, 51.59 W, 59 +2026/08/02 13:00:23.323, 63 %, 13964 MiB, 54.31 W, 59 +2026/08/02 13:00:23.825, 86 %, 13964 MiB, 54.31 W, 59 +2026/08/02 13:00:24.326, 85 %, 13964 MiB, 54.15 W, 59 +2026/08/02 13:00:24.828, 48 %, 13964 MiB, 51.37 W, 59 +2026/08/02 13:00:25.330, 99 %, 14056 MiB, 143.01 W, 61 +2026/08/02 13:00:25.831, 99 %, 14056 MiB, 144.13 W, 62 +2026/08/02 13:00:26.333, 99 %, 14056 MiB, 130.68 W, 63 +2026/08/02 13:00:26.834, 99 %, 14056 MiB, 149.54 W, 63 +2026/08/02 13:00:27.336, 94 %, 208 MiB, 53.70 W, 61 +2026/08/02 13:00:27.838, 0 %, 208 MiB, 51.94 W, 60 +2026/08/02 13:00:28.339, 0 %, 208 MiB, 51.83 W, 60 +2026/08/02 13:00:28.841, 0 %, 208 MiB, 51.77 W, 60 +2026/08/02 13:00:29.342, 79 %, 13964 MiB, 54.44 W, 59 +2026/08/02 13:00:29.844, 86 %, 13964 MiB, 54.28 W, 59 +2026/08/02 13:00:30.346, 86 %, 13964 MiB, 54.15 W, 59 +2026/08/02 13:00:30.849, 33 %, 13964 MiB, 51.51 W, 59 +2026/08/02 13:00:31.351, 99 %, 14156 MiB, 134.88 W, 61 +2026/08/02 13:00:31.854, 98 %, 14156 MiB, 135.21 W, 62 +2026/08/02 13:00:32.356, 99 %, 14156 MiB, 136.11 W, 63 +2026/08/02 13:00:32.858, 92 %, 14156 MiB, 138.08 W, 64 +2026/08/02 13:00:33.363, 99 %, 14156 MiB, 130.37 W, 64 +2026/08/02 13:00:33.865, 99 %, 14156 MiB, 135.03 W, 65 +2026/08/02 13:00:34.367, 99 %, 14156 MiB, 144.91 W, 65 +2026/08/02 13:00:34.868, 99 %, 14156 MiB, 139.84 W, 66 +2026/08/02 13:00:35.370, 99 %, 14156 MiB, 137.82 W, 66 +2026/08/02 13:00:35.872, 99 %, 14156 MiB, 139.90 W, 66 +2026/08/02 13:00:36.373, 99 %, 14156 MiB, 139.43 W, 66 +2026/08/02 13:00:36.875, 99 %, 14156 MiB, 144.74 W, 66 +2026/08/02 13:00:37.376, 99 %, 14156 MiB, 131.98 W, 66 +2026/08/02 13:00:37.879, 99 %, 14156 MiB, 136.58 W, 66 +2026/08/02 13:00:38.383, 99 %, 14156 MiB, 132.86 W, 66 +2026/08/02 13:00:38.886, 90 %, 208 MiB, 61.18 W, 64 +2026/08/02 13:00:39.388, 0 %, 208 MiB, 53.24 W, 63 diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-bench.jsonl b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-bench.jsonl new file mode 100644 index 000000000000..420537aaaa52 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-bench.jsonl @@ -0,0 +1,3 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1785674994,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-rFXXVTOCeizmTtKdbrsOL62vWx72JIFT","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":218.519,"prompt_per_token_ms":6.427029411764706,"prompt_per_second":155.59287750721904,"predicted_n":128,"predicted_ms":5323.693,"predicted_per_token_ms":41.5913515625,"predicted_per_second":24.04346005676886}} +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1785675000,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-hcfj6UnLHjOFc7kcSg1W4KTaxgiLbpCJ","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":210.052,"prompt_per_token_ms":6.178,"prompt_per_second":161.8646811265782,"predicted_n":128,"predicted_ms":5330.966,"predicted_per_token_ms":41.648171875,"predicted_per_second":24.010657730700213}} +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1785675005,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-DklyrNQ4SCFFuljBFuhrMPzE72g9ZJA1","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":210.056,"prompt_per_token_ms":6.178117647058824,"prompt_per_second":161.8615988117454,"predicted_n":128,"predicted_ms":5331.456,"predicted_per_token_ms":41.652,"predicted_per_second":24.00845097474311}} diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-smoke.json b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-smoke.json new file mode 100644 index 000000000000..0ea5326af161 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-smoke.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1785674665,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-k4vTvs6OHvXyLHI7S6WKh5lQqzTFAoDa","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":223.974,"prompt_per_token_ms":6.587470588235294,"prompt_per_second":151.8033343155902,"predicted_n":128,"predicted_ms":5293.112,"predicted_per_token_ms":41.3524375,"predicted_per_second":24.182371353562893}} \ No newline at end of file diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-telemetry.csv b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-telemetry.csv new file mode 100644 index 000000000000..85a0ddb339c0 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-baseline-telemetry.csv @@ -0,0 +1,34 @@ +2026/08/02 14:49:48.894, 52, 58.91, 0, 7520 +2026/08/02 14:49:49.396, 55, 138.92, 96, 7524 +2026/08/02 14:49:49.898, 56, 134.37, 98, 7524 +2026/08/02 14:49:50.399, 57, 144.93, 98, 7524 +2026/08/02 14:49:50.901, 58, 141.57, 98, 7524 +2026/08/02 14:49:51.402, 58, 145.34, 98, 7524 +2026/08/02 14:49:51.904, 59, 144.19, 98, 7524 +2026/08/02 14:49:52.405, 59, 144.87, 98, 7524 +2026/08/02 14:49:52.907, 59, 142.91, 98, 7524 +2026/08/02 14:49:53.409, 60, 142.44, 98, 7524 +2026/08/02 14:49:53.911, 60, 146.43, 98, 7524 +2026/08/02 14:49:54.413, 60, 134.83, 98, 7524 +2026/08/02 14:49:54.914, 60, 142.66, 97, 7524 +2026/08/02 14:49:55.416, 61, 141.26, 98, 7524 +2026/08/02 14:49:55.921, 61, 142.13, 98, 7524 +2026/08/02 14:49:56.424, 61, 147.30, 98, 7524 +2026/08/02 14:49:56.927, 61, 145.75, 98, 7524 +2026/08/02 14:49:57.428, 61, 143.63, 98, 7524 +2026/08/02 14:49:57.930, 62, 142.29, 98, 7524 +2026/08/02 14:49:58.431, 62, 135.46, 98, 7524 +2026/08/02 14:49:58.933, 62, 133.59, 98, 7524 +2026/08/02 14:49:59.434, 62, 133.90, 98, 7524 +2026/08/02 14:49:59.937, 62, 135.71, 98, 7524 +2026/08/02 14:50:00.440, 62, 146.33, 98, 7524 +2026/08/02 14:50:00.943, 62, 143.73, 98, 7524 +2026/08/02 14:50:01.445, 63, 149.10, 98, 7524 +2026/08/02 14:50:01.951, 63, 144.04, 98, 7524 +2026/08/02 14:50:02.452, 63, 145.12, 98, 7524 +2026/08/02 14:50:02.954, 63, 145.59, 98, 7524 +2026/08/02 14:50:03.457, 63, 143.82, 98, 7524 +2026/08/02 14:50:03.959, 63, 145.59, 98, 7524 +2026/08/02 14:50:04.462, 64, 145.63, 98, 7524 +2026/08/02 14:50:04.968, 64, 143.67, 98, 7524 +2026/08/02 14:50:05.472, 64, 138.82, 98, 7524 diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-bench.jsonl b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-bench.jsonl new file mode 100644 index 000000000000..d4ba302048dc --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-bench.jsonl @@ -0,0 +1,3 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674888,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-nQ8XKXSNLgA9Fx8SxoJouuaqpFgHXBv1","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":230.916,"prompt_per_token_ms":6.791647058823529,"prompt_per_second":147.23968889119854,"predicted_n":128,"predicted_ms":2264.98,"predicted_per_token_ms":17.69515625,"predicted_per_second":56.51264028821446,"draft_n":111,"draft_n_accepted":92}} +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674891,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-NPvBDfDHSvPeSINVFvJ1LxlZFUOgkcTX","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":211.042,"prompt_per_token_ms":6.207117647058824,"prompt_per_second":161.10537239032988,"predicted_n":128,"predicted_ms":2266.448,"predicted_per_token_ms":17.706625,"predicted_per_second":56.476036511757606,"draft_n":111,"draft_n_accepted":92}} +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674893,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-ZUy2pUU7xjyg0mKRiyvZIAB9saJ3DDnt","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":212.33,"prompt_per_token_ms":6.245,"prompt_per_second":160.1281024819856,"predicted_n":128,"predicted_ms":2271.916,"predicted_per_token_ms":17.74934375,"predicted_per_second":56.34011116608184,"draft_n":111,"draft_n_accepted":92}} diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-1.json b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-1.json new file mode 100644 index 000000000000..db7b45db628e --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674792,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-DpHjflfD3pRnc0PtHVfSqViWVQvfyOVc","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":222.55,"prompt_per_token_ms":6.545588235294118,"prompt_per_second":152.77465738036395,"predicted_n":128,"predicted_ms":2313.883,"predicted_per_token_ms":18.0772109375,"predicted_per_second":55.31826803688865,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-2.json b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-2.json new file mode 100644 index 000000000000..af471aa9cf7d --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-cpu-sampler-2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674794,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-bxG9BQ11U3WbqP6Mxt5YDJ12gPRx6tHx","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":211.351,"prompt_per_token_ms":6.216205882352941,"prompt_per_second":160.86983264805937,"predicted_n":128,"predicted_ms":2269.022,"predicted_per_token_ms":17.726734375,"predicted_per_second":56.411969562216676,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke-2.json b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke-2.json new file mode 100644 index 000000000000..acc34339bdb4 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke-2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1785674746,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-Uw5V0s5JuRxIAMaE2qW6zqBMrJfTL1Vs","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":176.032,"prompt_per_token_ms":6.519703703703704,"prompt_per_second":153.38120341756044,"predicted_n":128,"predicted_ms":2378.209,"predicted_per_token_ms":18.5797578125,"predicted_per_second":53.822014801895044,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke.json b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke.json new file mode 100644 index 000000000000..988d4f76b0c5 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-smoke.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1785674711,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-5pD48Aw2Bqat1v6hCzx8sMfeK09ye21W","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":239.541,"prompt_per_token_ms":7.045323529411765,"prompt_per_second":141.93812332753058,"predicted_n":128,"predicted_ms":2290.344,"predicted_per_token_ms":17.8933125,"predicted_per_second":55.88680128399926,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-telemetry.csv b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-telemetry.csv new file mode 100644 index 000000000000..1113136d5f48 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/gemma12-mtp-telemetry.csv @@ -0,0 +1,16 @@ +timestamp, utilization.gpu [%], memory.used [MiB], power.draw [W], temperature.gpu +2026/08/02 14:48:06.134, 0 %, 7824 MiB, 86.51 W, 46 +2026/08/02 14:48:06.635, 86 %, 7824 MiB, 158.87 W, 49 +2026/08/02 14:48:07.137, 93 %, 7824 MiB, 155.44 W, 51 +2026/08/02 14:48:07.639, 94 %, 7824 MiB, 123.99 W, 52 +2026/08/02 14:48:08.142, 91 %, 7824 MiB, 150.04 W, 53 +2026/08/02 14:48:08.644, 85 %, 7824 MiB, 172.38 W, 54 +2026/08/02 14:48:09.148, 90 %, 7824 MiB, 147.65 W, 55 +2026/08/02 14:48:09.652, 91 %, 7824 MiB, 90.35 W, 55 +2026/08/02 14:48:10.154, 94 %, 7824 MiB, 140.95 W, 55 +2026/08/02 14:48:10.655, 89 %, 7824 MiB, 131.20 W, 56 +2026/08/02 14:48:11.157, 88 %, 7824 MiB, 153.54 W, 56 +2026/08/02 14:48:11.658, 91 %, 7824 MiB, 142.64 W, 57 +2026/08/02 14:48:12.163, 93 %, 7824 MiB, 140.42 W, 57 +2026/08/02 14:48:12.669, 92 %, 7824 MiB, 112.58 W, 57 +2026/08/02 14:48:13.173, 93 %, 7824 MiB, 112.85 W, 57 diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.jsonl b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.jsonl new file mode 100644 index 000000000000..5a7856c23b06 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.jsonl @@ -0,0 +1,4 @@ +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/Qwen3.6-35B-A3B-Q4_K_M.gguf", "model_type": "qwen35moe 35B.A3B Q4_K - Medium", "model_size": 20408576512, "model_n_params": 34660610688, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 128, "n_gen": 0, "n_depth": 0, "test_time": "2026-08-02T10:58:08Z", "avg_ns": 569115973, "stddev_ns": 42618002, "avg_ts": 225.722631, "stddev_ts": 16.276246, "samples_ns": [ 617792032, 538511005, 551044883 ],"samples_ts": [ 207.189, 237.692, 232.286 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/Qwen3.6-35B-A3B-Q4_K_M.gguf", "model_type": "qwen35moe 35B.A3B Q4_K - Medium", "model_size": 20408576512, "model_n_params": 34660610688, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 512, "n_gen": 0, "n_depth": 0, "test_time": "2026-08-02T10:58:32Z", "avg_ns": 1177401787, "stddev_ns": 238997360, "avg_ts": 445.845497, "stddev_ts": 81.236366, "samples_ns": [ 1452793028, 1024230838, 1055181496 ],"samples_ts": [ 352.425, 499.887, 485.225 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/Qwen3.6-35B-A3B-Q4_K_M.gguf", "model_type": "qwen35moe 35B.A3B Q4_K - Medium", "model_size": 20408576512, "model_n_params": 34660610688, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 0, "n_gen": 32, "n_depth": 0, "test_time": "2026-08-02T10:59:02Z", "avg_ns": 1068717619, "stddev_ns": 158549904, "avg_ts": 30.353757, "stddev_ts": 4.159909, "samples_ns": [ 1251215401, 990077854, 964859604 ],"samples_ts": [ 25.5751, 32.3207, 33.1654 ]} +{"build_commit": "798cf6cbe", "build_number": 9862, "cpu_info": "Intel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz", "gpu_info": "Quadro P5200", "backends": "CUDA", "model_filename": "/home/me/Models/Qwen3.6-35B-A3B-Q4_K_M.gguf", "model_type": "qwen35moe 35B.A3B Q4_K - Medium", "model_size": 20408576512, "model_n_params": 34660610688, "n_batch": 2048, "n_ubatch": 512, "n_threads": 4, "cpu_mask": "0x0", "cpu_strict": false, "poll": 50, "type_k": "f16", "type_v": "f16", "n_gpu_layers": -1, "n_cpu_moe": 0, "split_mode": "layer", "main_gpu": 0, "no_kv_offload": false, "flash_attn": 1, "devices": "auto", "tensor_split": "0.00", "tensor_buft_overrides": "none", "use_mmap": true, "use_direct_io": false, "embeddings": false, "no_op_offload": 0, "no_host": false, "fit_target": 1024, "fit_min_ctx": 4096, "n_prompt": 0, "n_gen": 128, "n_depth": 0, "test_time": "2026-08-02T10:59:26Z", "avg_ns": 5363713708, "stddev_ns": 4073818412, "avg_ts": 27.504544, "stddev_ts": 10.792292, "samples_ns": [ 8498035317, 3864710510, 3728395297 ],"samples_ts": [ 15.0623, 33.1202, 34.3311 ]} diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.log b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.log new file mode 100644 index 000000000000..e24d4dfa87f3 --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-bench.log @@ -0,0 +1,22 @@ +ggml_cuda_init: found 1 CUDA devices (Total VRAM: 16266 MiB): + Device 0: Quadro P5200, compute capability 6.1, VMM: yes, VRAM: 16266 MiB +llama-bench: benchmark 1/4: starting +llama-bench: benchmark 1/4: warmup prompt run +llama-bench: benchmark 1/4: prompt run 1/3 +llama-bench: benchmark 1/4: prompt run 2/3 +llama-bench: benchmark 1/4: prompt run 3/3 +llama-bench: benchmark 2/4: starting +llama-bench: benchmark 2/4: warmup prompt run +llama-bench: benchmark 2/4: prompt run 1/3 +llama-bench: benchmark 2/4: prompt run 2/3 +llama-bench: benchmark 2/4: prompt run 3/3 +llama-bench: benchmark 3/4: starting +llama-bench: benchmark 3/4: warmup generation run +llama-bench: benchmark 3/4: generation run 1/3 +llama-bench: benchmark 3/4: generation run 2/3 +llama-bench: benchmark 3/4: generation run 3/3 +llama-bench: benchmark 4/4: starting +llama-bench: benchmark 4/4: warmup generation run +llama-bench: benchmark 4/4: generation run 1/3 +llama-bench: benchmark 4/4: generation run 2/3 +llama-bench: benchmark 4/4: generation run 3/3 diff --git a/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-telemetry.csv b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-telemetry.csv new file mode 100644 index 000000000000..6356f546d2ff --- /dev/null +++ b/docs/benchmarks/data/pascal-p5200-2026-08-02/qwen-telemetry.csv @@ -0,0 +1,239 @@ +timestamp, utilization.gpu [%], memory.used [MiB], power.draw [W], temperature.gpu +2026/08/02 12:57:45.509, 0 %, 90 MiB, 9.23 W, 48 +2026/08/02 12:57:46.019, 0 %, 198 MiB, 39.13 W, 48 +2026/08/02 12:57:46.520, 0 %, 262 MiB, 39.30 W, 49 +2026/08/02 12:57:47.022, 1 %, 198 MiB, 39.28 W, 49 +2026/08/02 12:57:47.524, 0 %, 198 MiB, 39.30 W, 49 +2026/08/02 12:57:48.025, 0 %, 198 MiB, 39.57 W, 49 +2026/08/02 12:57:48.527, 0 %, 198 MiB, 39.38 W, 49 +2026/08/02 12:57:49.028, 0 %, 198 MiB, 39.40 W, 49 +2026/08/02 12:57:49.530, 0 %, 198 MiB, 39.40 W, 50 +2026/08/02 12:57:50.032, 0 %, 198 MiB, 39.38 W, 50 +2026/08/02 12:57:50.533, 0 %, 198 MiB, 39.53 W, 50 +2026/08/02 12:57:51.035, 0 %, 198 MiB, 39.57 W, 50 +2026/08/02 12:57:51.537, 0 %, 198 MiB, 39.37 W, 50 +2026/08/02 12:57:52.038, 0 %, 198 MiB, 39.38 W, 50 +2026/08/02 12:57:52.540, 0 %, 198 MiB, 39.43 W, 50 +2026/08/02 12:57:53.041, 0 %, 198 MiB, 39.54 W, 50 +2026/08/02 12:57:53.543, 0 %, 198 MiB, 39.54 W, 50 +2026/08/02 12:57:54.047, 0 %, 198 MiB, 39.54 W, 50 +2026/08/02 12:57:54.549, 0 %, 198 MiB, 39.43 W, 50 +2026/08/02 12:57:55.051, 0 %, 198 MiB, 39.46 W, 50 +2026/08/02 12:57:55.552, 0 %, 198 MiB, 39.58 W, 50 +2026/08/02 12:57:56.054, 0 %, 198 MiB, 39.54 W, 50 +2026/08/02 12:57:56.556, 0 %, 198 MiB, 39.37 W, 50 +2026/08/02 12:57:57.058, 0 %, 198 MiB, 39.57 W, 50 +2026/08/02 12:57:57.559, 0 %, 198 MiB, 39.57 W, 50 +2026/08/02 12:57:58.061, 0 %, 198 MiB, 39.58 W, 51 +2026/08/02 12:57:58.563, 0 %, 198 MiB, 39.56 W, 51 +2026/08/02 12:57:59.064, 0 %, 198 MiB, 39.54 W, 51 +2026/08/02 12:57:59.566, 0 %, 198 MiB, 39.53 W, 51 +2026/08/02 12:58:00.082, 0 %, 198 MiB, 39.57 W, 51 +2026/08/02 12:58:00.583, 0 %, 198 MiB, 39.57 W, 51 +2026/08/02 12:58:01.085, 21 %, 14778 MiB, 39.83 W, 51 +2026/08/02 12:58:01.586, 20 %, 14778 MiB, 40.03 W, 51 +2026/08/02 12:58:02.088, 23 %, 14778 MiB, 40.00 W, 51 +2026/08/02 12:58:02.590, 15 %, 14778 MiB, 39.87 W, 51 +2026/08/02 12:58:03.091, 20 %, 14778 MiB, 39.86 W, 51 +2026/08/02 12:58:03.593, 20 %, 14778 MiB, 40.01 W, 51 +2026/08/02 12:58:04.095, 20 %, 14778 MiB, 39.98 W, 51 +2026/08/02 12:58:04.596, 19 %, 14778 MiB, 40.16 W, 51 +2026/08/02 12:58:05.098, 18 %, 14778 MiB, 40.00 W, 51 +2026/08/02 12:58:05.599, 17 %, 14778 MiB, 39.84 W, 51 +2026/08/02 12:58:06.101, 11 %, 14778 MiB, 40.01 W, 51 +2026/08/02 12:58:06.603, 7 %, 14778 MiB, 39.57 W, 51 +2026/08/02 12:58:07.104, 19 %, 14778 MiB, 40.00 W, 51 +2026/08/02 12:58:07.606, 21 %, 14778 MiB, 40.16 W, 51 +2026/08/02 12:58:08.108, 27 %, 14778 MiB, 40.34 W, 51 +2026/08/02 12:58:08.609, 12 %, 14848 MiB, 40.36 W, 51 +2026/08/02 12:58:09.111, 19 %, 15036 MiB, 120.97 W, 52 +2026/08/02 12:58:09.612, 37 %, 15036 MiB, 42.99 W, 52 +2026/08/02 12:58:10.114, 17 %, 15036 MiB, 42.85 W, 52 +2026/08/02 12:58:10.616, 91 %, 15036 MiB, 87.66 W, 55 +2026/08/02 12:58:11.117, 93 %, 15036 MiB, 142.60 W, 56 +2026/08/02 12:58:11.623, 90 %, 15036 MiB, 138.56 W, 57 +2026/08/02 12:58:12.125, 79 %, 208 MiB, 59.59 W, 55 +2026/08/02 12:58:12.626, 0 %, 208 MiB, 51.51 W, 54 +2026/08/02 12:58:13.128, 0 %, 208 MiB, 51.51 W, 54 +2026/08/02 12:58:13.630, 0 %, 272 MiB, 51.35 W, 54 +2026/08/02 12:58:14.131, 0 %, 208 MiB, 51.20 W, 54 +2026/08/02 12:58:14.633, 0 %, 208 MiB, 51.28 W, 54 +2026/08/02 12:58:15.134, 0 %, 208 MiB, 51.12 W, 54 +2026/08/02 12:58:15.636, 0 %, 208 MiB, 51.28 W, 54 +2026/08/02 12:58:16.138, 0 %, 208 MiB, 51.10 W, 54 +2026/08/02 12:58:16.639, 0 %, 208 MiB, 51.16 W, 54 +2026/08/02 12:58:17.141, 0 %, 208 MiB, 51.14 W, 54 +2026/08/02 12:58:17.645, 0 %, 208 MiB, 51.18 W, 54 +2026/08/02 12:58:18.147, 0 %, 208 MiB, 39.87 W, 53 +2026/08/02 12:58:18.651, 0 %, 208 MiB, 39.94 W, 53 +2026/08/02 12:58:19.193, 0 %, 208 MiB, 39.84 W, 52 +2026/08/02 12:58:19.695, 0 %, 208 MiB, 39.86 W, 53 +2026/08/02 12:58:20.197, 0 %, 208 MiB, 39.84 W, 52 +2026/08/02 12:58:20.799, 0 %, 208 MiB, 39.73 W, 52 +2026/08/02 12:58:21.301, 0 %, 208 MiB, 39.73 W, 52 +2026/08/02 12:58:21.803, 0 %, 208 MiB, 39.86 W, 52 +2026/08/02 12:58:22.304, 0 %, 208 MiB, 39.68 W, 52 +2026/08/02 12:58:22.806, 0 %, 208 MiB, 39.71 W, 52 +2026/08/02 12:58:23.308, 0 %, 208 MiB, 39.86 W, 52 +2026/08/02 12:58:23.809, 18 %, 14356 MiB, 40.30 W, 52 +2026/08/02 12:58:24.311, 20 %, 14356 MiB, 40.31 W, 52 +2026/08/02 12:58:24.813, 22 %, 14356 MiB, 40.33 W, 52 +2026/08/02 12:58:25.314, 24 %, 14356 MiB, 40.06 W, 52 +2026/08/02 12:58:25.816, 20 %, 14356 MiB, 39.73 W, 52 +2026/08/02 12:58:26.318, 20 %, 14356 MiB, 40.16 W, 52 +2026/08/02 12:58:26.822, 3 %, 14356 MiB, 39.84 W, 52 +2026/08/02 12:58:27.544, 0 %, 14356 MiB, 39.71 W, 52 +2026/08/02 12:58:28.046, 17 %, 14356 MiB, 40.30 W, 52 +2026/08/02 12:58:28.549, 21 %, 14356 MiB, 40.19 W, 52 +2026/08/02 12:58:29.051, 14 %, 14356 MiB, 39.71 W, 52 +2026/08/02 12:58:29.553, 2 %, 14356 MiB, 39.71 W, 52 +2026/08/02 12:58:30.054, 20 %, 14356 MiB, 39.71 W, 52 +2026/08/02 12:58:30.556, 19 %, 14356 MiB, 40.16 W, 52 +2026/08/02 12:58:31.058, 20 %, 14356 MiB, 40.16 W, 52 +2026/08/02 12:58:31.559, 20 %, 14356 MiB, 40.17 W, 52 +2026/08/02 12:58:32.061, 22 %, 14356 MiB, 40.31 W, 52 +2026/08/02 12:58:32.563, 10 %, 14356 MiB, 39.86 W, 52 +2026/08/02 12:58:33.064, 0 %, 14928 MiB, 39.71 W, 52 +2026/08/02 12:58:33.567, 100 %, 14956 MiB, 144.41 W, 56 +2026/08/02 12:58:34.068, 10 %, 14956 MiB, 48.32 W, 54 +2026/08/02 12:58:34.570, 5 %, 14956 MiB, 48.19 W, 54 +2026/08/02 12:58:35.071, 13 %, 14956 MiB, 48.76 W, 54 +2026/08/02 12:58:35.574, 12 %, 14956 MiB, 50.65 W, 54 +2026/08/02 12:58:36.075, 11 %, 14956 MiB, 146.15 W, 54 +2026/08/02 12:58:36.577, 11 %, 14956 MiB, 118.99 W, 54 +2026/08/02 12:58:37.081, 8 %, 14956 MiB, 132.34 W, 56 +2026/08/02 12:58:37.586, 78 %, 14956 MiB, 52.83 W, 57 +2026/08/02 12:58:38.088, 75 %, 14956 MiB, 161.82 W, 57 +2026/08/02 12:58:38.589, 100 %, 14956 MiB, 141.82 W, 59 +2026/08/02 12:58:39.094, 80 %, 14956 MiB, 83.84 W, 58 +2026/08/02 12:58:39.596, 100 %, 14956 MiB, 147.36 W, 60 +2026/08/02 12:58:40.098, 78 %, 14956 MiB, 104.00 W, 59 +2026/08/02 12:58:40.600, 96 %, 208 MiB, 60.26 W, 58 +2026/08/02 12:58:41.101, 0 %, 208 MiB, 52.12 W, 57 +2026/08/02 12:58:41.603, 0 %, 272 MiB, 51.98 W, 56 +2026/08/02 12:58:42.104, 1 %, 272 MiB, 51.83 W, 56 +2026/08/02 12:58:42.606, 0 %, 208 MiB, 51.83 W, 56 +2026/08/02 12:58:43.107, 0 %, 208 MiB, 51.73 W, 56 +2026/08/02 12:58:43.609, 0 %, 208 MiB, 51.67 W, 55 +2026/08/02 12:58:44.110, 0 %, 208 MiB, 51.59 W, 55 +2026/08/02 12:58:44.612, 0 %, 208 MiB, 51.71 W, 55 +2026/08/02 12:58:45.120, 0 %, 208 MiB, 51.59 W, 55 +2026/08/02 12:58:45.622, 0 %, 208 MiB, 51.73 W, 55 +2026/08/02 12:58:46.123, 0 %, 208 MiB, 51.45 W, 55 +2026/08/02 12:58:46.625, 0 %, 208 MiB, 51.57 W, 55 +2026/08/02 12:58:47.129, 0 %, 208 MiB, 51.59 W, 55 +2026/08/02 12:58:47.630, 0 %, 208 MiB, 51.57 W, 55 +2026/08/02 12:58:48.132, 0 %, 208 MiB, 51.61 W, 55 +2026/08/02 12:58:48.633, 0 %, 208 MiB, 51.57 W, 55 +2026/08/02 12:58:49.137, 0 %, 208 MiB, 51.43 W, 55 +2026/08/02 12:58:49.639, 0 %, 208 MiB, 51.63 W, 55 +2026/08/02 12:58:50.141, 0 %, 208 MiB, 51.61 W, 55 +2026/08/02 12:58:50.656, 0 %, 208 MiB, 51.59 W, 55 +2026/08/02 12:58:51.165, 0 %, 208 MiB, 51.55 W, 55 +2026/08/02 12:58:51.667, 0 %, 208 MiB, 51.57 W, 55 +2026/08/02 12:58:52.168, 0 %, 208 MiB, 51.55 W, 55 +2026/08/02 12:58:52.670, 0 %, 208 MiB, 51.39 W, 55 +2026/08/02 12:58:53.172, 2 %, 14788 MiB, 52.20 W, 55 +2026/08/02 12:58:53.673, 13 %, 14788 MiB, 52.18 W, 55 +2026/08/02 12:58:54.175, 21 %, 14788 MiB, 52.20 W, 55 +2026/08/02 12:58:54.677, 21 %, 14788 MiB, 52.49 W, 55 +2026/08/02 12:58:55.178, 19 %, 14788 MiB, 52.04 W, 55 +2026/08/02 12:58:55.680, 20 %, 14788 MiB, 52.36 W, 55 +2026/08/02 12:58:56.181, 22 %, 14788 MiB, 52.18 W, 55 +2026/08/02 12:58:56.683, 22 %, 14788 MiB, 52.18 W, 55 +2026/08/02 12:58:57.185, 20 %, 14788 MiB, 51.77 W, 55 +2026/08/02 12:58:57.686, 2 %, 14788 MiB, 51.63 W, 55 +2026/08/02 12:58:58.188, 1 %, 14788 MiB, 52.22 W, 55 +2026/08/02 12:58:58.692, 14 %, 14788 MiB, 51.90 W, 55 +2026/08/02 12:58:59.194, 16 %, 14788 MiB, 52.04 W, 55 +2026/08/02 12:58:59.697, 14 %, 14788 MiB, 51.73 W, 55 +2026/08/02 12:59:00.200, 16 %, 14788 MiB, 52.20 W, 55 +2026/08/02 12:59:00.702, 15 %, 14788 MiB, 52.02 W, 55 +2026/08/02 12:59:01.205, 18 %, 14788 MiB, 52.32 W, 55 +2026/08/02 12:59:01.709, 17 %, 14788 MiB, 52.32 W, 55 +2026/08/02 12:59:02.211, 9 %, 15008 MiB, 57.81 W, 55 +2026/08/02 12:59:02.713, 36 %, 15008 MiB, 110.26 W, 57 +2026/08/02 12:59:03.215, 60 %, 15008 MiB, 88.21 W, 58 +2026/08/02 12:59:03.716, 59 %, 15008 MiB, 138.53 W, 59 +2026/08/02 12:59:04.218, 70 %, 15008 MiB, 110.20 W, 60 +2026/08/02 12:59:04.720, 69 %, 15008 MiB, 133.84 W, 60 +2026/08/02 12:59:05.222, 71 %, 15008 MiB, 120.78 W, 61 +2026/08/02 12:59:05.724, 100 %, 208 MiB, 60.42 W, 59 +2026/08/02 12:59:06.225, 0 %, 208 MiB, 52.46 W, 58 +2026/08/02 12:59:06.727, 0 %, 272 MiB, 52.30 W, 58 +2026/08/02 12:59:07.229, 1 %, 272 MiB, 52.12 W, 57 +2026/08/02 12:59:07.731, 0 %, 208 MiB, 52.14 W, 57 +2026/08/02 12:59:08.232, 0 %, 208 MiB, 52.08 W, 57 +2026/08/02 12:59:08.734, 0 %, 208 MiB, 40.47 W, 56 +2026/08/02 12:59:09.236, 0 %, 208 MiB, 40.46 W, 56 +2026/08/02 12:59:09.737, 0 %, 208 MiB, 40.33 W, 56 +2026/08/02 12:59:10.239, 0 %, 208 MiB, 40.34 W, 55 +2026/08/02 12:59:10.741, 0 %, 208 MiB, 40.34 W, 55 +2026/08/02 12:59:11.242, 0 %, 208 MiB, 40.30 W, 55 +2026/08/02 12:59:11.744, 0 %, 208 MiB, 40.20 W, 55 +2026/08/02 12:59:12.245, 0 %, 208 MiB, 40.14 W, 55 +2026/08/02 12:59:12.750, 0 %, 208 MiB, 40.16 W, 55 +2026/08/02 12:59:13.251, 0 %, 208 MiB, 40.28 W, 55 +2026/08/02 12:59:13.753, 0 %, 208 MiB, 40.20 W, 55 +2026/08/02 12:59:14.254, 0 %, 208 MiB, 40.20 W, 55 +2026/08/02 12:59:14.756, 0 %, 208 MiB, 40.25 W, 55 +2026/08/02 12:59:15.260, 0 %, 208 MiB, 40.17 W, 55 +2026/08/02 12:59:15.762, 0 %, 208 MiB, 40.17 W, 55 +2026/08/02 12:59:16.264, 0 %, 208 MiB, 40.17 W, 55 +2026/08/02 12:59:16.765, 0 %, 208 MiB, 40.13 W, 54 +2026/08/02 12:59:17.267, 0 %, 208 MiB, 40.14 W, 55 +2026/08/02 12:59:17.769, 0 %, 208 MiB, 40.16 W, 54 +2026/08/02 12:59:18.270, 20 %, 14642 MiB, 40.47 W, 54 +2026/08/02 12:59:18.772, 20 %, 14642 MiB, 40.64 W, 54 +2026/08/02 12:59:19.273, 20 %, 14642 MiB, 40.49 W, 54 +2026/08/02 12:59:19.775, 19 %, 14642 MiB, 40.64 W, 54 +2026/08/02 12:59:20.277, 5 %, 14642 MiB, 40.31 W, 54 +2026/08/02 12:59:20.780, 0 %, 14642 MiB, 40.17 W, 54 +2026/08/02 12:59:21.328, 4 %, 14642 MiB, 40.20 W, 55 +2026/08/02 12:59:21.833, 10 %, 14642 MiB, 40.47 W, 54 +2026/08/02 12:59:22.334, 21 %, 14642 MiB, 40.94 W, 54 +2026/08/02 12:59:22.836, 28 %, 14642 MiB, 40.33 W, 54 +2026/08/02 12:59:23.337, 20 %, 14642 MiB, 40.61 W, 54 +2026/08/02 12:59:23.839, 20 %, 14642 MiB, 40.61 W, 54 +2026/08/02 12:59:24.340, 21 %, 14642 MiB, 41.13 W, 54 +2026/08/02 12:59:24.842, 49 %, 14642 MiB, 53.20 W, 55 +2026/08/02 12:59:25.344, 1 %, 14642 MiB, 51.61 W, 55 +2026/08/02 12:59:25.851, 7 %, 14642 MiB, 51.79 W, 55 +2026/08/02 12:59:26.359, 0 %, 14642 MiB, 51.65 W, 56 +2026/08/02 12:59:26.861, 9 %, 14872 MiB, 51.71 W, 56 +2026/08/02 12:59:27.362, 0 %, 14872 MiB, 51.79 W, 56 +2026/08/02 12:59:27.864, 1 %, 14872 MiB, 52.26 W, 56 +2026/08/02 12:59:28.366, 10 %, 14872 MiB, 54.11 W, 56 +2026/08/02 12:59:28.867, 11 %, 14872 MiB, 52.26 W, 56 +2026/08/02 12:59:29.369, 1 %, 14872 MiB, 53.63 W, 56 +2026/08/02 12:59:29.871, 1 %, 14872 MiB, 52.06 W, 56 +2026/08/02 12:59:30.372, 1 %, 14872 MiB, 52.49 W, 56 +2026/08/02 12:59:30.874, 1 %, 14872 MiB, 53.16 W, 56 +2026/08/02 12:59:31.376, 11 %, 14872 MiB, 53.01 W, 57 +2026/08/02 12:59:31.878, 33 %, 14872 MiB, 69.61 W, 58 +2026/08/02 12:59:32.380, 39 %, 14872 MiB, 87.97 W, 58 +2026/08/02 12:59:32.882, 49 %, 14872 MiB, 79.93 W, 59 +2026/08/02 12:59:33.383, 43 %, 14872 MiB, 74.81 W, 60 +2026/08/02 12:59:33.885, 39 %, 14872 MiB, 114.33 W, 59 +2026/08/02 12:59:34.387, 52 %, 14872 MiB, 100.83 W, 60 +2026/08/02 12:59:34.888, 65 %, 14872 MiB, 116.51 W, 61 +2026/08/02 12:59:35.390, 69 %, 14872 MiB, 130.06 W, 61 +2026/08/02 12:59:35.892, 67 %, 14872 MiB, 111.31 W, 62 +2026/08/02 12:59:36.393, 50 %, 14872 MiB, 81.99 W, 61 +2026/08/02 12:59:36.895, 63 %, 14872 MiB, 106.25 W, 62 +2026/08/02 12:59:37.396, 67 %, 14872 MiB, 93.31 W, 62 +2026/08/02 12:59:37.898, 68 %, 14872 MiB, 131.46 W, 62 +2026/08/02 12:59:38.400, 73 %, 14872 MiB, 113.18 W, 62 +2026/08/02 12:59:38.901, 72 %, 14872 MiB, 120.57 W, 63 +2026/08/02 12:59:39.403, 68 %, 14872 MiB, 129.85 W, 63 +2026/08/02 12:59:39.904, 66 %, 14872 MiB, 134.20 W, 63 +2026/08/02 12:59:40.406, 70 %, 14872 MiB, 93.89 W, 63 +2026/08/02 12:59:40.907, 67 %, 14872 MiB, 135.13 W, 63 +2026/08/02 12:59:41.409, 73 %, 14872 MiB, 140.97 W, 64 +2026/08/02 12:59:41.911, 72 %, 14872 MiB, 131.87 W, 64 +2026/08/02 12:59:42.413, 73 %, 14872 MiB, 142.12 W, 64 +2026/08/02 12:59:42.914, 71 %, 14872 MiB, 104.76 W, 64 +2026/08/02 12:59:43.416, 71 %, 14872 MiB, 118.44 W, 64 +2026/08/02 12:59:43.921, 86 %, 208 MiB, 60.24 W, 62 +2026/08/02 12:59:44.423, 0 %, 208 MiB, 52.57 W, 61 +2026/08/02 12:59:44.924, 0 %, 208 MiB, 51.98 W, 61 diff --git a/docs/benchmarks/pascal-p5200-mtp.md b/docs/benchmarks/pascal-p5200-mtp.md new file mode 100644 index 000000000000..8d7ee38fa317 --- /dev/null +++ b/docs/benchmarks/pascal-p5200-mtp.md @@ -0,0 +1,123 @@ +# MTP speculative decoding on Pascal (Quadro P5200, sm_61) + +Measured results for `--spec-type draft-mtp` on a compute-capability 6.1 GPU, +plus two Pascal-specific behaviours that are easy to hit and are not obvious +from the option list. + +All numbers below were recomputed from the raw per-request timings in +[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/), not copied +from a summary. + +## Test setup + +| | | +| --- | --- | +| GPU | NVIDIA Quadro P5200 Mobile, 16 GiB, compute capability **6.1** | +| CPU | Intel Core i7-7820HQ, 4C/8T | +| Build | `b9862-798cf6cbe`, `CMAKE_CUDA_ARCHITECTURES=61`, `GGML_CUDA_FORCE_MMQ=ON`, `GGML_CUDA_F16=OFF` | +| Target | `gemma-4-12b-it-qat-q4_0.gguf` — 6,975,879,296 B, `93567e57a8fe10b2…` | +| Draft | `mtp-gemma-4-12B-it-Q4_0.gguf` — 253,708,960 B, `b894e614824dfc27…` | +| Server | one slot, 4096-token context, full offload, flash attention, greedy sampling | +| Samples | three independent 128-token requests per mode | + +Both GGUFs were exact-size and SHA-256 verified before testing. The target is +`google/gemma-4-12B-it-qat-q4_0-gguf`; the assistant is the QAT-derived +`gemma4-assistant` GGUF from `ggml-org/gemma-4-12B-it-GGUF`. + +## Results + +| Mode | Generation | Prompt processing | Peak VRAM | Peak GPU | Peak power | +| --- | ---: | ---: | ---: | ---: | ---: | +| Target only | 24.02 tok/s | 159.77 tok/s | 7,524 MiB | 98% | 149.10 W | +| MTP | **56.44 tok/s** | 156.16 tok/s | 7,824 MiB | 94% | 172.38 W | + +**2.35x generation throughput** for **+300 MiB VRAM** and **+23 W**. Prompt +processing is unchanged within sample spread. + +Per-sample generation rates were tight in both modes — baseline 24.04 / 24.01 / +24.01 tok/s, MTP 56.51 / 56.48 / 56.34 tok/s — so the ratio is not an artifact +of one fast run. Draft acceptance was **276 of 333 proposed tokens (82.88%)**, +i.e. 92 of 111 in each of the three runs. + +On a 16 GiB card where VRAM is the binding constraint, +300 MiB for 2.35x is a +favourable trade; the assistant model itself is only 254 MB on disk. + +## Pascal caution 1 — `-md` alone does not enable MTP + +Supplying a draft model without an explicit `--spec-type draft-mtp` does not +turn MTP on. The server starts, loads both models and serves normally, so the +failure mode is silent: you get baseline throughput and no error. + +Confirm from the response `timings`: if `draft_n` is `0`, MTP is not running. + +## Pascal caution 2 — draft backend sampling is non-repeatable on sm_61 + +`--spec-draft-backend-sampling` defaults to **enabled** +(`common/common.h`: `backend_sampling = true`, "offload draft sampling to the +backend"). With it enabled on this P5200, repeated **greedy** requests with +identical input produced **different** outputs. + +Passing `--no-spec-draft-backend-sampling` made repeated MTP runs byte-identical +to one another without materially reducing speed — the 56.44 tok/s above was +measured with backend sampling disabled. + +This is reported as an observation on one sm_61 device, not as a diagnosis. It +has not been bisected, and no claim is made about other architectures. The +isolation runs are in `gemma12-mtp-cpu-sampler-1.json` and +`gemma12-mtp-cpu-sampler-2.json`. + +Note that the flag is currently absent from the option list in +[`../speculative.md`](../speculative.md); it is added there by the same change +that adds this page. + +## What "deterministic" does and does not mean here + +With `--no-spec-draft-backend-sampling`, all three MTP responses were +byte-identical **to one another**, and all three target-only responses were +byte-identical to one another. + +The MTP response was **not** byte-identical to the target-only response. This is +expected rather than a defect: the target still verifies every proposed token, +but batched target evaluation can select a different greedy token than +single-token evaluation because GPU floating-point reduction order differs. + +So this result demonstrates verified-token speculation and run-to-run +repeatability. It should **not** be described as bit-identical to +non-speculative generation. + +## Reproducing + +```sh +./build-cuda/bin/llama serve \ + -m /path/to/gemma-4-12b-it-qat-q4_0.gguf \ + -md /path/to/mtp-gemma-4-12B-it-Q4_0.gguf \ + -c 4096 -ngl all -ngld all -fa on --parallel 1 \ + --spec-type draft-mtp --spec-draft-n-max 16 --spec-draft-p-min 0.9 \ + --no-spec-draft-backend-sampling \ + --host 127.0.0.1 --port 8080 --no-webui --jinja +``` + +Gemma's embedded chat template requires `--jinja`; without it the legacy +template path reports that the custom template is unsupported. + +## Scope + +Single device, single model pair, `--parallel 1`. Speculative gains normally +shrink as concurrency rises, because the target batch fills with real work, so +**2.35x is an upper bound for a multi-slot server**, not a general figure. No +multi-slot measurement was taken. + +## Raw data + +[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/) — verbatim, as +produced by the run: + +- `README.md` — the original run report, including a same-build `llama-bench` + pass over Qwen3.6-35B-A3B Q4_K_M and Gemma 4 26B-A4B QAT Q4_0 on the same GPU +- `gemma12-baseline-bench.jsonl`, `gemma12-mtp-bench.jsonl` — three API + responses per mode with llama.cpp timings and MTP acceptance counters +- `gemma12-baseline-telemetry.csv`, `gemma12-mtp-telemetry.csv` — 500 ms + `nvidia-smi` samples +- `gemma12-*-smoke.json`, `gemma12-mtp-cpu-sampler-*.json` — smoke and + determinism-isolation runs +- `qwen-*`, `gemma-*` — the `llama-bench` pass and its telemetry diff --git a/docs/speculative.md b/docs/speculative.md index 8f91256c4a4d..57b4f8a48802 100644 --- a/docs/speculative.md +++ b/docs/speculative.md @@ -176,6 +176,9 @@ If a draft model is combined with a draftless decoding the draftless decoding ha --spec-draft-p-min, --draft-p-min P minimum speculative decoding probability (greedy) (default: 0.00) (env: LLAMA_ARG_SPEC_DRAFT_P_MIN) +--spec-draft-backend-sampling, --no-spec-draft-backend-sampling + offload draft sampling to the backend (default: enabled) + (env: LLAMA_ARG_SPEC_DRAFT_BACKEND_SAMPLING) --spec-draft-ngl, -ngld, --gpu-layers-draft, --n-gpu-layers-draft N max. number of draft model layers to store in VRAM, either an exact number, 'auto', or 'all' (default: auto) (env: LLAMA_ARG_N_GPU_LAYERS_DRAFT) @@ -368,3 +371,5 @@ statistics ngram_map_k: #calls(b,g,a) = 6 1690 26, #gen drafts = 26, #acc drafts To measure the end-to-end effect of speculative decoding (throughput, latency, and draft acceptance) across diverse prompts, see the SPEED-Bench client in [tools/server/bench/speed-bench](../tools/server/bench/speed-bench/README.md). It runs against a running `llama-server` and can compare a baseline run against a speculative-decoding run. + +For a measured `draft-mtp` baseline-vs-speculative comparison on a Pascal (sm_61) GPU, including two hardware-specific cautions, see [benchmarks/pascal-p5200-mtp.md](benchmarks/pascal-p5200-mtp.md). From d5f459e87d35c0f7255e2c9470251c5fd7cc755d Mon Sep 17 00:00:00 2001 From: Agent Date: Fri, 7 Aug 2026 12:47:29 +0200 Subject: [PATCH 2/2] docs(speculative): correct both Pascal MTP cautions after reproducing them Reproduced both cautions from scratch on the same P5200 rather than carrying them over from the run report. Both were wrong, in opposite directions, and the corrected versions are stronger. Caution 2 was refuted. It claimed draft backend sampling produced non-repeatable greedy output on sm_61. Six conditions, five identical greedy requests each: A MTP, draft backend sampling enabled (default) 2 distinct / 5 B MTP, draft backend sampling disabled 2 distinct / 5 C target only, no draft model (control) 2 distinct / 5 D target only, -bs 2 distinct / 5 E target only, cache_prompt=false 1 distinct / 5 F MTP, cache_prompt=false 1 distinct / 5 A and B are byte-identical to each other, so the flag changes nothing. The no-draft control diverges identically, so MTP is not involved. The actual cause is prompt cache reuse changing the prompt batch split (1st request cache_n=0 prompt_n=34; later requests cache_n=7 prompt_n=27), which changes reduction order and can flip a greedy token. `cache_prompt: false` restores repeatability in both conditions. Nothing here is sm_61-specific. Caution 1 understated the failure. It said `-md` without an explicit `--spec-type draft-mtp` silently yields baseline throughput. In fact the server auto-enables `draft-simple`, which an MTP head model cannot satisfy; startup completes, `GET /health` reports healthy, and every completion request then fails with HTTP 500 "decode() failed: failed to process speculative batch". The preceding "failed to create llama_context from model" line is logged at warning level and is not treated as fatal. Adds the repro scripts and all response artifacts under docs/benchmarks/data/repro-2026-08-07/ so both claims can be re-checked. Throughput numbers are unchanged and were already recomputed from the raw timings. Acting agent: crystal-mom --- .../benchmarks/data/repro-2026-08-07/A.1.json | 1 + .../benchmarks/data/repro-2026-08-07/A.2.json | 1 + .../benchmarks/data/repro-2026-08-07/A.3.json | 1 + .../benchmarks/data/repro-2026-08-07/A.4.json | 1 + .../benchmarks/data/repro-2026-08-07/A.5.json | 1 + .../benchmarks/data/repro-2026-08-07/B.1.json | 1 + .../benchmarks/data/repro-2026-08-07/B.2.json | 1 + .../benchmarks/data/repro-2026-08-07/B.3.json | 1 + .../benchmarks/data/repro-2026-08-07/B.4.json | 1 + .../benchmarks/data/repro-2026-08-07/B.5.json | 1 + .../benchmarks/data/repro-2026-08-07/C.1.json | 1 + .../benchmarks/data/repro-2026-08-07/C.2.json | 1 + .../benchmarks/data/repro-2026-08-07/C.3.json | 1 + .../benchmarks/data/repro-2026-08-07/C.4.json | 1 + .../benchmarks/data/repro-2026-08-07/C.5.json | 1 + .../benchmarks/data/repro-2026-08-07/D.1.json | 1 + .../benchmarks/data/repro-2026-08-07/D.2.json | 1 + .../benchmarks/data/repro-2026-08-07/D.3.json | 1 + .../benchmarks/data/repro-2026-08-07/D.4.json | 1 + .../benchmarks/data/repro-2026-08-07/D.5.json | 1 + .../benchmarks/data/repro-2026-08-07/E.1.json | 1 + .../benchmarks/data/repro-2026-08-07/E.2.json | 1 + .../benchmarks/data/repro-2026-08-07/E.3.json | 1 + .../benchmarks/data/repro-2026-08-07/E.4.json | 1 + .../benchmarks/data/repro-2026-08-07/E.5.json | 1 + .../benchmarks/data/repro-2026-08-07/F.1.json | 1 + .../benchmarks/data/repro-2026-08-07/F.2.json | 1 + .../benchmarks/data/repro-2026-08-07/F.3.json | 1 + .../benchmarks/data/repro-2026-08-07/F.4.json | 1 + .../benchmarks/data/repro-2026-08-07/F.5.json | 1 + .../data/repro-2026-08-07/G2.1.json | 1 + .../repro-2026-08-07/G2.server.log.excerpt | 12 ++ .../data/repro-2026-08-07/mtp-determinism.sh | 109 +++++++++++++++ .../data/repro-2026-08-07/mtp-determinism2.sh | 87 ++++++++++++ .../data/repro-2026-08-07/mtp-g2.sh | 18 +++ .../data/repro-2026-08-07/summary.txt | 6 + docs/benchmarks/pascal-p5200-mtp.md | 126 +++++++++++++----- 37 files changed, 355 insertions(+), 34 deletions(-) create mode 100644 docs/benchmarks/data/repro-2026-08-07/A.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/A.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/A.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/A.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/A.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/B.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/B.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/B.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/B.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/B.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/C.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/C.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/C.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/C.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/C.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/D.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/D.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/D.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/D.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/D.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/E.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/E.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/E.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/E.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/E.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/F.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/F.2.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/F.3.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/F.4.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/F.5.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/G2.1.json create mode 100644 docs/benchmarks/data/repro-2026-08-07/G2.server.log.excerpt create mode 100755 docs/benchmarks/data/repro-2026-08-07/mtp-determinism.sh create mode 100755 docs/benchmarks/data/repro-2026-08-07/mtp-determinism2.sh create mode 100755 docs/benchmarks/data/repro-2026-08-07/mtp-g2.sh create mode 100644 docs/benchmarks/data/repro-2026-08-07/summary.txt diff --git a/docs/benchmarks/data/repro-2026-08-07/A.1.json b/docs/benchmarks/data/repro-2026-08-07/A.1.json new file mode 100644 index 000000000000..8a5c9554a862 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/A.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099239,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-1qmgdpY37ZVb91U8ZC2gkipd2XRsHmZm","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":251.688,"prompt_per_token_ms":7.402588235294117,"prompt_per_second":135.08788658974603,"predicted_n":128,"predicted_ms":2515.941,"predicted_per_token_ms":19.6557890625,"predicted_per_second":50.87559684428212,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/A.2.json b/docs/benchmarks/data/repro-2026-08-07/A.2.json new file mode 100644 index 000000000000..7fc152f9631e --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/A.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099241,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-Bw8rovjgQeFIOIdmgQSRphJGrVamGIb8","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":190.566,"prompt_per_token_ms":7.058,"prompt_per_second":141.68319637291017,"predicted_n":128,"predicted_ms":2615.732,"predicted_per_token_ms":20.43540625,"predicted_per_second":48.93467679410582,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/A.3.json b/docs/benchmarks/data/repro-2026-08-07/A.3.json new file mode 100644 index 000000000000..2b7dda7dadb1 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/A.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099244,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-QxyxrfYw183l7zXBT94C2aX3hy84e9QY","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":190.481,"prompt_per_token_ms":7.0548518518518515,"prompt_per_second":141.74642090287222,"predicted_n":128,"predicted_ms":2603.991,"predicted_per_token_ms":20.3436796875,"predicted_per_second":49.15531582098402,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/A.4.json b/docs/benchmarks/data/repro-2026-08-07/A.4.json new file mode 100644 index 000000000000..c7017ab4b0d0 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/A.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099247,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-d6MzEo5A6o42M0EAkROUUCYpq0VRW5Iw","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.619,"prompt_per_token_ms":7.097,"prompt_per_second":140.90460758066789,"predicted_n":128,"predicted_ms":2577.429,"predicted_per_token_ms":20.1361640625,"predicted_per_second":49.661891753371286,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/A.5.json b/docs/benchmarks/data/repro-2026-08-07/A.5.json new file mode 100644 index 000000000000..418849f3237f --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/A.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099250,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-lx6y1oPU8VQHuGI4P1INpvj53CJxk8au","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.717,"prompt_per_token_ms":7.10062962962963,"prompt_per_second":140.83258135689582,"predicted_n":128,"predicted_ms":2595.115,"predicted_per_token_ms":20.2743359375,"predicted_per_second":49.32344038703488,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/B.1.json b/docs/benchmarks/data/repro-2026-08-07/B.1.json new file mode 100644 index 000000000000..586570eaafa4 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/B.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099263,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-WIO4nbIcukqCIdTwOSa9YDgg0U3ggLQl","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":242.11,"prompt_per_token_ms":7.120882352941177,"prompt_per_second":140.43203502540166,"predicted_n":128,"predicted_ms":2567.172,"predicted_per_token_ms":20.05603125,"predicted_per_second":49.860313216255086,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/B.2.json b/docs/benchmarks/data/repro-2026-08-07/B.2.json new file mode 100644 index 000000000000..a4b575569e14 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/B.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099266,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-YKOY8ALeMRazVaO1O4ncPWN6w0e9Ir4D","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":190.232,"prompt_per_token_ms":7.045629629629629,"prompt_per_second":141.93195676857732,"predicted_n":128,"predicted_ms":2667.37,"predicted_per_token_ms":20.838828125,"predicted_per_second":47.987343338194556,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/B.3.json b/docs/benchmarks/data/repro-2026-08-07/B.3.json new file mode 100644 index 000000000000..66a93ee7d6b0 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/B.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099269,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-azHp0s31ExQS4qaSp3IanjsL8fLrZr2Y","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":196.134,"prompt_per_token_ms":7.264222222222222,"prompt_per_second":137.66098687631927,"predicted_n":128,"predicted_ms":2690.257,"predicted_per_token_ms":21.0176328125,"predicted_per_second":47.57909746169232,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/B.4.json b/docs/benchmarks/data/repro-2026-08-07/B.4.json new file mode 100644 index 000000000000..5ccf445bd14d --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/B.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099272,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-0IqFmUAf9EZbRRf1O8ObjFDlYWIP7FK7","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.297,"prompt_per_token_ms":7.085074074074074,"prompt_per_second":141.1417847640057,"predicted_n":128,"predicted_ms":2702.647,"predicted_per_token_ms":21.1144296875,"predicted_per_second":47.36097610971762,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/B.5.json b/docs/benchmarks/data/repro-2026-08-07/B.5.json new file mode 100644 index 000000000000..89e4a66614eb --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/B.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition of RAID (Redundancy vs. Backup). RAID protects against hardware failure, not data loss.\n * Paragraph 2: The \"Human/Software Error\" factor. Deleting a file or a virus affects all drives in the array simultaneously.\n * Paragraph 3: The \"Disaster Recovery\" factor. Fire, theft, or natural disasters destroy the physical hardware"}}],"created":1786099275,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-9libthPnHAy1y7vhw1PeQQOlGYSVh7cu","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.351,"prompt_per_token_ms":7.087074074074074,"prompt_per_second":141.10195400076302,"predicted_n":128,"predicted_ms":2667.791,"predicted_per_token_ms":20.8421171875,"predicted_per_second":47.9797705292506,"draft_n":116,"draft_n_accepted":88}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/C.1.json b/docs/benchmarks/data/repro-2026-08-07/C.1.json new file mode 100644 index 000000000000..5854313b5b1c --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/C.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099290,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-DXv2d0aglGrHG9cMStiNptbVVP92lFpv","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":231.31,"prompt_per_token_ms":6.8032352941176475,"prompt_per_second":146.98888936924476,"predicted_n":128,"predicted_ms":5811.406,"predicted_per_token_ms":45.401609375,"predicted_per_second":22.025650935419073}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/C.2.json b/docs/benchmarks/data/repro-2026-08-07/C.2.json new file mode 100644 index 000000000000..32d0b2d882e8 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/C.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099296,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-mQvyOvRqU0T6KPDpjmJDKEYpQusNIsc8","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":188.47,"prompt_per_token_ms":6.980370370370371,"prompt_per_second":143.25887409136735,"predicted_n":128,"predicted_ms":5825.019,"predicted_per_token_ms":45.5079609375,"predicted_per_second":21.97417725161068}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/C.3.json b/docs/benchmarks/data/repro-2026-08-07/C.3.json new file mode 100644 index 000000000000..4367bd0d4b6c --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/C.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099302,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-jQUycJq5XkIBiGEdXewgR32jDB2oMLMa","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.883,"prompt_per_token_ms":7.106777777777778,"prompt_per_second":140.71074561060647,"predicted_n":128,"predicted_ms":5831.497,"predicted_per_token_ms":45.5585703125,"predicted_per_second":21.949766929486543}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/C.4.json b/docs/benchmarks/data/repro-2026-08-07/C.4.json new file mode 100644 index 000000000000..561bd474b08e --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/C.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099308,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-VWk9KYqZNYmXNSJUrQ5cQoL0HpOCro4c","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":193.671,"prompt_per_token_ms":7.173,"prompt_per_second":139.41168269901019,"predicted_n":128,"predicted_ms":5820.365,"predicted_per_token_ms":45.4716015625,"predicted_per_second":21.9917479402065}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/C.5.json b/docs/benchmarks/data/repro-2026-08-07/C.5.json new file mode 100644 index 000000000000..8ead9aebf80e --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/C.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099314,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-xI4sj3MVhLbeP8HKlTdEhOniMEQr9pkT","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":191.921,"prompt_per_token_ms":7.108185185185185,"prompt_per_second":140.68288514545048,"predicted_n":128,"predicted_ms":5838.104,"predicted_per_token_ms":45.6101875,"predicted_per_second":21.924926311692975}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/D.1.json b/docs/benchmarks/data/repro-2026-08-07/D.1.json new file mode 100644 index 000000000000..d76d5623dea4 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/D.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099329,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-xgeFg0Pg6GYfFtgvZJOC4hflsccECXX4","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":250.632,"prompt_per_token_ms":7.371529411764706,"prompt_per_second":135.65705895496185,"predicted_n":128,"predicted_ms":5770.141,"predicted_per_token_ms":45.0792265625,"predicted_per_second":22.183166754503922}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/D.2.json b/docs/benchmarks/data/repro-2026-08-07/D.2.json new file mode 100644 index 000000000000..8e4c959ed530 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/D.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099335,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-6AYYTXroe4Ek2e4r0NjCDxBTw3oMUrNL","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":220.099,"prompt_per_token_ms":8.151814814814815,"prompt_per_second":122.6720702956397,"predicted_n":128,"predicted_ms":5778.602,"predicted_per_token_ms":45.145328125,"predicted_per_second":22.150686273254326}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/D.3.json b/docs/benchmarks/data/repro-2026-08-07/D.3.json new file mode 100644 index 000000000000..6cb0255c4f24 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/D.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099341,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-dmyBDLdErkDUfbRnxmqVkDhnC2cqGEEd","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":234.777,"prompt_per_token_ms":8.695444444444444,"prompt_per_second":115.00274728785188,"predicted_n":128,"predicted_ms":5772.706,"predicted_per_token_ms":45.099265625,"predicted_per_second":22.173310055977215}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/D.4.json b/docs/benchmarks/data/repro-2026-08-07/D.4.json new file mode 100644 index 000000000000..05255c7fd1c8 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/D.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099347,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-omCPuXCmKlzUCssHVBgyqZIKoGwsiKbd","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":227.628,"prompt_per_token_ms":8.430666666666665,"prompt_per_second":118.6145816859086,"predicted_n":128,"predicted_ms":5774.017,"predicted_per_token_ms":45.1095078125,"predicted_per_second":22.168275569677057}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/D.5.json b/docs/benchmarks/data/repro-2026-08-07/D.5.json new file mode 100644 index 000000000000..fc8abc63134f --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/D.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Data Integrity/Human Error (Deletion, Corruption).\n * Paragraph 3: Hardware/System Failure (Ransomware, Physical Disaster).\n * Paragraph 4: Practical Example and Conclusion.\n\n * *Paragraph 1:* RAID (Redundant Array of Independent"}}],"created":1786099353,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":7}},"id":"chatcmpl-5OoVFPBRQwaXrm9uhOIsKW9zRqLK2QlK","timings":{"cache_n":7,"prompt_n":27,"prompt_ms":213.252,"prompt_per_token_ms":7.8982222222222225,"prompt_per_second":126.6107703561983,"predicted_n":128,"predicted_ms":5776.05,"predicted_per_token_ms":45.125390625,"predicted_per_second":22.160472987595327}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/E.1.json b/docs/benchmarks/data/repro-2026-08-07/E.1.json new file mode 100644 index 000000000000..62dc1b0fba6b --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/E.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099428,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-fWx0bOYT8tHkcC6FeAeZNy9fTB17YQfr","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":233.075,"prompt_per_token_ms":6.855147058823529,"prompt_per_second":145.87579105438164,"predicted_n":128,"predicted_ms":5822.57,"predicted_per_token_ms":45.488828125,"predicted_per_second":21.983419692678662}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/E.2.json b/docs/benchmarks/data/repro-2026-08-07/E.2.json new file mode 100644 index 000000000000..8d7ecb4a8042 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/E.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099435,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-W9f3qhS6ERqmBxBxoEFNglWRkZwKchNs","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":226.52,"prompt_per_token_ms":6.662352941176471,"prompt_per_second":150.09712166696096,"predicted_n":128,"predicted_ms":5823.13,"predicted_per_token_ms":45.493203125,"predicted_per_second":21.98130558651447}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/E.3.json b/docs/benchmarks/data/repro-2026-08-07/E.3.json new file mode 100644 index 000000000000..f672636044aa --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/E.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099441,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-Kekm6MQaN1L872w6vzvrpm5Jx2oHz1ts","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":225.941,"prompt_per_token_ms":6.6453235294117645,"prompt_per_second":150.48176293811215,"predicted_n":128,"predicted_ms":5815.002,"predicted_per_token_ms":45.429703125,"predicted_per_second":22.012030262414353}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/E.4.json b/docs/benchmarks/data/repro-2026-08-07/E.4.json new file mode 100644 index 000000000000..872bb459da62 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/E.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099447,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-0a2fyQKWfD16895IIK6qINRqqStEnnpb","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":227.502,"prompt_per_token_ms":6.691235294117647,"prompt_per_second":149.4492356111155,"predicted_n":128,"predicted_ms":5823.788,"predicted_per_token_ms":45.49834375,"predicted_per_second":21.97882203129647}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/E.5.json b/docs/benchmarks/data/repro-2026-08-07/E.5.json new file mode 100644 index 000000000000..c7018eaee70e --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/E.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * Paragraph 1: Definition/Purpose of RAID (Redundancy vs. Backup).\n * Paragraph 2: Failure modes (Human error, malware, hardware failure).\n * Paragraph 3: The practical example (Deleting a file).\n * Paragraph 4: Conclusion/Best practice (3-2-1 rule).\n\n * *Paragraph 1:* RAID (Red"}}],"created":1786099453,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-YsQ0DZZC4ihsS3ldJFoB8zJMT30lWccI","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":224.868,"prompt_per_token_ms":6.613764705882353,"prompt_per_second":151.1998150025793,"predicted_n":128,"predicted_ms":5841.067,"predicted_per_token_ms":45.6333359375,"predicted_per_second":21.913804447029968}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/F.1.json b/docs/benchmarks/data/repro-2026-08-07/F.1.json new file mode 100644 index 000000000000..789a868a1813 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/F.1.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099466,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-QFOHmpo2k8nNEFhDf3Sgxakexu53CIgz","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":244.07,"prompt_per_token_ms":7.178529411764706,"prompt_per_second":139.3042979473102,"predicted_n":128,"predicted_ms":2533.818,"predicted_per_token_ms":19.795453125,"predicted_per_second":50.516651156476115,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/F.2.json b/docs/benchmarks/data/repro-2026-08-07/F.2.json new file mode 100644 index 000000000000..f3c70c3e28f0 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/F.2.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099469,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-Z4nhuaVoY9RqZuoyKoPtxninU42XyeTG","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":229.047,"prompt_per_token_ms":6.736676470588235,"prompt_per_second":148.44114963304474,"predicted_n":128,"predicted_ms":2465.709,"predicted_per_token_ms":19.2633515625,"predicted_per_second":51.912046393146966,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/F.3.json b/docs/benchmarks/data/repro-2026-08-07/F.3.json new file mode 100644 index 000000000000..c9051126472f --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/F.3.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099472,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-n8W6fi2U6bv4HFV8XlmaGMWKFGdXQzny","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":234.281,"prompt_per_token_ms":6.890617647058824,"prompt_per_second":145.12487141509553,"predicted_n":128,"predicted_ms":2467.039,"predicted_per_token_ms":19.2737421875,"predicted_per_second":51.88406020334498,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/F.4.json b/docs/benchmarks/data/repro-2026-08-07/F.4.json new file mode 100644 index 000000000000..d702b53e0ce9 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/F.4.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099474,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-7KPhdLfj6xJYT8MDXNtbBTBvw3l7tJ59","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":230.667,"prompt_per_token_ms":6.784323529411765,"prompt_per_second":147.3986309268339,"predicted_n":128,"predicted_ms":2470.738,"predicted_per_token_ms":19.302640625,"predicted_per_second":51.80638335590419,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/F.5.json b/docs/benchmarks/data/repro-2026-08-07/F.5.json new file mode 100644 index 000000000000..fe1849df2a48 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/F.5.json @@ -0,0 +1 @@ +{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"","reasoning_content":"* Topic: Why RAID is not a backup.\n * Constraint 1: Exactly four short paragraphs.\n * Constraint 2: Include one practical example.\n\n * *Paragraph 1: Definition/Purpose of RAID.* RAID (Redundant Array of Independent Disks) is designed for hardware redundancy and performance. It protects against hardware failure (like a single drive dying) but doesn't protect against data loss from other sources.\n * *Paragraph 2: Data Integrity/Human Error.* RAID doesn't protect against accidental deletion, file corruption, or malware. If"}}],"created":1786099477,"model":"/home/me/Models/gemma-4-12b-it-qat-q4_0.gguf","system_fingerprint":"b9862-798cf6cbe","object":"chat.completion","usage":{"completion_tokens":128,"prompt_tokens":34,"total_tokens":162,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-HMN8JWh6uCS76sdX6DbJtu8NJKlGJYwj","timings":{"cache_n":0,"prompt_n":34,"prompt_ms":235.033,"prompt_per_token_ms":6.912735294117646,"prompt_per_second":144.66053703097012,"predicted_n":128,"predicted_ms":2455.544,"predicted_per_token_ms":19.1839375,"predicted_per_second":52.126942135836295,"draft_n":111,"draft_n_accepted":92}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/G2.1.json b/docs/benchmarks/data/repro-2026-08-07/G2.1.json new file mode 100644 index 000000000000..d522056f663c --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/G2.1.json @@ -0,0 +1 @@ +{"error":{"code":500,"message":"decode() failed: failed to process speculative batch","type":"server_error"}} \ No newline at end of file diff --git a/docs/benchmarks/data/repro-2026-08-07/G2.server.log.excerpt b/docs/benchmarks/data/repro-2026-08-07/G2.server.log.excerpt new file mode 100644 index 000000000000..97c562beb0ce --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/G2.server.log.excerpt @@ -0,0 +1,12 @@ +0.00.914.909 E llama_init_from_model: failed to initialize the context: Gemma4Assistant requires ctx_other to be set (this warning is normal during memory fitting) +0.01.012.700 W srv load_model: [spec] failed to measure draft model memory: failed to create llama_context from model +0.02.787.723 W load: special_eog_ids contains '<|tool_response>', removing '' token from EOG list +0.03.669.099 W load: special_eog_ids contains '<|tool_response>', removing '' token from EOG list +0.06.297.997 W load: special_eog_ids contains '<|tool_response>', removing '' token from EOG list +0.06.459.503 W common_speculative_init: draft model is specified but 'draft' speculative type is not explicitly enabled - enabling it +0.06.459.547 I common_speculative_impl_draft_simple: adding speculative implementation 'draft-simple' +0.06.459.549 I common_speculative_impl_draft_simple: - n_max=3, n_min=0, p_min=0.000000 +0.06.459.552 I common_speculative_impl_draft_simple: - gpu_layers=-2, cache_k=f16, cache_v=f16, ctx_tgt=yes, ctx_dft=yes, devices=[default] +0.06.466.766 I srv load_model: speculative decoding context initialized +0.07.206.856 E decode: failed to initialize batch +0.07.206.857 E llama_decode: failed to decode, ret = -1 diff --git a/docs/benchmarks/data/repro-2026-08-07/mtp-determinism.sh b/docs/benchmarks/data/repro-2026-08-07/mtp-determinism.sh new file mode 100755 index 000000000000..3cf320bec372 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/mtp-determinism.sh @@ -0,0 +1,109 @@ +#!/usr/bin/env bash +# Minimal repro: is greedy output repeatable on sm_61 with backend sampling? +# +# Four conditions, same prompt, same seed, greedy, N identical requests each. +# A MTP, draft backend sampling ENABLED (shipped default) +# B MTP, draft backend sampling DISABLED +# C target only, main backend sampling DISABLED (shipped default) [control] +# D target only, main backend sampling ENABLED (-bs) [discriminator] +# +# C is the control: if C diverges, the box/build is non-deterministic for +# reasons unrelated to backend sampling and the whole experiment is void. +# D discriminates "draft path is broken" from "backend sampling is broken". + +set -uo pipefail + +BIN="$HOME/ht/ht-llama.cpp/build-cuda/bin/llama" +TGT="$HOME/Models/gemma-4-12b-it-qat-q4_0.gguf" +DRAFT="$HOME/Models/mtp-gemma-4-12B-it-Q4_0.gguf" +PORT=8099 +N=5 +OUT="$HOME/mtp-determinism" +PROMPT="Explain in exactly four short paragraphs why RAID is not a backup. Include one practical example." + +rm -rf "$OUT"; mkdir -p "$OUT" + +echo "build: $("$BIN" --version 2>&1 | head -1)" +echo "gpu: $(nvidia-smi --query-gpu=name,driver_version --format=csv,noheader)" +echo "target: $(sha256sum "$TGT" | cut -c1-16) draft: $(sha256sum "$DRAFT" | cut -c1-16)" +echo + +start_server() { + # shellcheck disable=SC2068 + "$BIN" serve $@ \ + -c 4096 -ngl all -fa on --parallel 1 \ + --host 127.0.0.1 --port "$PORT" --no-webui --jinja \ + > "$OUT/$COND.server.log" 2>&1 & + SRV=$! + for _ in $(seq 1 300); do + if curl -sf -m 2 "http://127.0.0.1:$PORT/health" >/dev/null 2>&1; then return 0; fi + if ! kill -0 "$SRV" 2>/dev/null; then echo " server died, see $OUT/$COND.server.log"; return 1; fi + sleep 1 + done + echo " server never became healthy"; return 1 +} + +stop_server() { + kill "$SRV" 2>/dev/null + wait "$SRV" 2>/dev/null + sleep 3 +} + +run_condition() { + COND=$1; shift + DESC=$1; shift + echo "=== $COND: $DESC" + if ! start_server "$@"; then echo " SKIPPED"; return 1; fi + + for i in $(seq 1 "$N"); do + curl -sf -m 300 "http://127.0.0.1:$PORT/v1/chat/completions" \ + -H 'Content-Type: application/json' \ + -d "{\"messages\":[{\"role\":\"user\",\"content\":$(printf '%s' "$PROMPT" | python3 -c 'import json,sys;print(json.dumps(sys.stdin.read()))')}], + \"max_tokens\":128,\"temperature\":0,\"top_k\":1,\"top_p\":1,\"seed\":42,\"stream\":false}" \ + > "$OUT/$COND.$i.json" + done + stop_server + + python3 - "$COND" "$OUT" "$N" <<'PY' +import hashlib, json, sys, os +cond, out, n = sys.argv[1], sys.argv[2], int(sys.argv[3]) +digests, drafts = [], [] +for i in range(1, n+1): + p = os.path.join(out, f"{cond}.{i}.json") + try: + o = json.load(open(p)) + except Exception as e: + print(f" run {i}: UNREADABLE ({e})"); digests.append(f"err{i}"); continue + m = o["choices"][0]["message"] + text = (m.get("content") or "") + (m.get("reasoning_content") or "") + t = o.get("timings", {}) + digests.append(hashlib.sha256(text.encode()).hexdigest()[:16]) + drafts.append((t.get("draft_n", 0), t.get("draft_n_accepted", 0))) + print(f" run {i}: sha={digests[-1]} len={len(text):4d} draft_n={t.get('draft_n',0)}") +uniq = len(set(digests)) +mtp_active = any(d[0] > 0 for d in drafts) +print(f" -> {uniq} distinct output(s) across {n} runs MTP active: {mtp_active}") +print(f" -> VERDICT: {'REPEATABLE' if uniq==1 else 'NON-REPEATABLE'}") +open(os.path.join(out, "summary.txt"), "a").write( + f"{cond}\tdistinct={uniq}/{n}\tmtp_active={mtp_active}\n") +PY + echo +} + +run_condition A "MTP, draft backend sampling ENABLED (default)" \ + -m "$TGT" -md "$DRAFT" -ngld all --spec-type draft-mtp \ + --spec-draft-n-max 16 --spec-draft-p-min 0.9 + +run_condition B "MTP, draft backend sampling DISABLED" \ + -m "$TGT" -md "$DRAFT" -ngld all --spec-type draft-mtp \ + --spec-draft-n-max 16 --spec-draft-p-min 0.9 \ + --no-spec-draft-backend-sampling + +run_condition C "target only, main backend sampling DISABLED (control)" \ + -m "$TGT" + +run_condition D "target only, main backend sampling ENABLED (-bs)" \ + -m "$TGT" -bs + +echo "=== SUMMARY ===" +cat "$OUT/summary.txt" diff --git a/docs/benchmarks/data/repro-2026-08-07/mtp-determinism2.sh b/docs/benchmarks/data/repro-2026-08-07/mtp-determinism2.sh new file mode 100755 index 000000000000..16dac608e31f --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/mtp-determinism2.sh @@ -0,0 +1,87 @@ +#!/usr/bin/env bash +# Follow-up: test the prompt-cache hypothesis. +# +# Hypothesis: the run-1-vs-rest divergence is caused by prompt cache reuse +# changing the prompt batch split (run 1: cache_n=0 prompt_n=34; +# runs 2+: cache_n=7 prompt_n=27), not by backend sampling or MTP. +# +# Prediction, which can fail: with "cache_prompt": false every request +# evaluates the full prompt identically, so all N runs must be identical -- +# in BOTH the target-only and the MTP condition. +# +# E target only, cache_prompt=false +# F MTP, cache_prompt=false +# +# If E or F still diverges, the hypothesis is wrong. + +set -uo pipefail + +BIN="$HOME/ht/ht-llama.cpp/build-cuda/bin/llama" +TGT="$HOME/Models/gemma-4-12b-it-qat-q4_0.gguf" +DRAFT="$HOME/Models/mtp-gemma-4-12B-it-Q4_0.gguf" +PORT=8099 +N=5 +OUT="$HOME/mtp-determinism" +PROMPT="Explain in exactly four short paragraphs why RAID is not a backup. Include one practical example." + +mkdir -p "$OUT" + +start_server() { + # shellcheck disable=SC2068 + "$BIN" serve $@ \ + -c 4096 -ngl all -fa on --parallel 1 \ + --host 127.0.0.1 --port "$PORT" --no-webui --jinja \ + > "$OUT/$COND.server.log" 2>&1 & + SRV=$! + for _ in $(seq 1 300); do + curl -sf -m 2 "http://127.0.0.1:$PORT/health" >/dev/null 2>&1 && return 0 + kill -0 "$SRV" 2>/dev/null || { echo " server died"; return 1; } + sleep 1 + done + echo " never healthy"; return 1 +} + +run_condition() { + COND=$1; shift + DESC=$1; shift + echo "=== $COND: $DESC" + start_server "$@" || { echo " SKIPPED"; return 1; } + + for i in $(seq 1 "$N"); do + curl -sf -m 300 "http://127.0.0.1:$PORT/v1/chat/completions" \ + -H 'Content-Type: application/json' \ + -d "{\"messages\":[{\"role\":\"user\",\"content\":$(printf '%s' "$PROMPT" | python3 -c 'import json,sys;print(json.dumps(sys.stdin.read()))')}], + \"max_tokens\":128,\"temperature\":0,\"top_k\":1,\"top_p\":1,\"seed\":42, + \"cache_prompt\":false,\"stream\":false}" \ + > "$OUT/$COND.$i.json" + done + kill "$SRV" 2>/dev/null; wait "$SRV" 2>/dev/null; sleep 3 + + python3 - "$COND" "$OUT" "$N" <<'PY' +import hashlib, json, sys, os +cond, out, n = sys.argv[1], sys.argv[2], int(sys.argv[3]) +digests = [] +for i in range(1, n+1): + o = json.load(open(os.path.join(out, f"{cond}.{i}.json"))) + m = o["choices"][0]["message"] + text = (m.get("content") or "") + (m.get("reasoning_content") or "") + t = o.get("timings", {}) + digests.append(hashlib.sha256(text.encode()).hexdigest()[:16]) + print(f" run {i}: sha={digests[-1]} len={len(text):4d} " + f"cache_n={t.get('cache_n')} prompt_n={t.get('prompt_n')} draft_n={t.get('draft_n')}") +uniq = len(set(digests)) +print(f" -> {uniq} distinct output(s) across {n} runs") +print(f" -> VERDICT: {'REPEATABLE' if uniq==1 else 'NON-REPEATABLE'}") +open(os.path.join(out, "summary.txt"), "a").write(f"{cond}\tdistinct={uniq}/{n}\n") +PY + echo +} + +run_condition E "target only, cache_prompt=false" -m "$TGT" + +run_condition F "MTP, cache_prompt=false" \ + -m "$TGT" -md "$DRAFT" -ngld all --spec-type draft-mtp \ + --spec-draft-n-max 16 --spec-draft-p-min 0.9 + +echo "=== SUMMARY (all conditions) ===" +cat "$OUT/summary.txt" diff --git a/docs/benchmarks/data/repro-2026-08-07/mtp-g2.sh b/docs/benchmarks/data/repro-2026-08-07/mtp-g2.sh new file mode 100755 index 000000000000..ef7cfba52116 --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/mtp-g2.sh @@ -0,0 +1,18 @@ +#!/usr/bin/env bash +set -uo pipefail +BIN="$HOME/ht/ht-llama.cpp/build-cuda/bin/llama" +TGT="$HOME/Models/gemma-4-12b-it-qat-q4_0.gguf" +DRAFT="$HOME/Models/mtp-gemma-4-12B-it-Q4_0.gguf" +PORT=8099; OUT="$HOME/mtp-determinism" +PROMPT="Explain in exactly four short paragraphs why RAID is not a backup. Include one practical example." +"$BIN" serve -m "$TGT" -md "$DRAFT" -ngld all \ + -c 4096 -ngl all -fa on --parallel 1 --host 127.0.0.1 --port $PORT --no-webui --jinja \ + > "$OUT/G2.server.log" 2>&1 & +SRV=$! +for _ in $(seq 1 300); do curl -sf -m 2 http://127.0.0.1:$PORT/health >/dev/null 2>&1 && break; kill -0 $SRV 2>/dev/null || { echo "SERVER DIED"; break; }; sleep 1; done +echo "server alive: $(kill -0 $SRV 2>/dev/null && echo yes || echo no)" +code=$(curl -s -o "$OUT/G2.1.json" -w "%{http_code}" -m 300 http://127.0.0.1:$PORT/v1/chat/completions -H "Content-Type: application/json" \ + -d "{\"messages\":[{\"role\":\"user\",\"content\":$(printf %s "$PROMPT" | python3 -c "import json,sys;print(json.dumps(sys.stdin.read()))")}],\"max_tokens\":128,\"temperature\":0,\"top_k\":1,\"top_p\":1,\"seed\":42,\"cache_prompt\":false,\"stream\":false}") +echo "http=$code bytes=$(stat -c%s "$OUT/G2.1.json")" +head -c 300 "$OUT/G2.1.json"; echo +kill $SRV 2>/dev/null; wait $SRV 2>/dev/null diff --git a/docs/benchmarks/data/repro-2026-08-07/summary.txt b/docs/benchmarks/data/repro-2026-08-07/summary.txt new file mode 100644 index 000000000000..2c17a3b8965a --- /dev/null +++ b/docs/benchmarks/data/repro-2026-08-07/summary.txt @@ -0,0 +1,6 @@ +A distinct=2/5 mtp_active=True +B distinct=2/5 mtp_active=True +C distinct=2/5 mtp_active=False +D distinct=2/5 mtp_active=False +E distinct=1/5 +F distinct=1/5 diff --git a/docs/benchmarks/pascal-p5200-mtp.md b/docs/benchmarks/pascal-p5200-mtp.md index 8d7ee38fa317..48e56f037c63 100644 --- a/docs/benchmarks/pascal-p5200-mtp.md +++ b/docs/benchmarks/pascal-p5200-mtp.md @@ -1,18 +1,20 @@ # MTP speculative decoding on Pascal (Quadro P5200, sm_61) Measured results for `--spec-type draft-mtp` on a compute-capability 6.1 GPU, -plus two Pascal-specific behaviours that are easy to hit and are not obvious -from the option list. +plus two behaviours that are easy to hit and are not obvious from the option +list. All numbers below were recomputed from the raw per-request timings in -[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/), not copied +[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/), and both +cautions were reproduced from scratch with the scripts in +[`data/repro-2026-08-07/`](data/repro-2026-08-07/) rather than carried over from a summary. ## Test setup | | | | --- | --- | -| GPU | NVIDIA Quadro P5200 Mobile, 16 GiB, compute capability **6.1** | +| GPU | NVIDIA Quadro P5200 Mobile, 16 GiB, compute capability **6.1**, driver 580.159.04 | | CPU | Intel Core i7-7820HQ, 4C/8T | | Build | `b9862-798cf6cbe`, `CMAKE_CUDA_ARCHITECTURES=61`, `GGML_CUDA_FORCE_MMQ=ON`, `GGML_CUDA_F16=OFF` | | Target | `gemma-4-12b-it-qat-q4_0.gguf` — 6,975,879,296 B, `93567e57a8fe10b2…` | @@ -42,44 +44,92 @@ i.e. 92 of 111 in each of the three runs. On a 16 GiB card where VRAM is the binding constraint, +300 MiB for 2.35x is a favourable trade; the assistant model itself is only 254 MB on disk. -## Pascal caution 1 — `-md` alone does not enable MTP +## Caution 1 — `-md` without `--spec-type` makes a healthy server that 500s -Supplying a draft model without an explicit `--spec-type draft-mtp` does not -turn MTP on. The server starts, loads both models and serves normally, so the -failure mode is silent: you get baseline throughput and no error. +Passing an MTP assistant via `-md` **without** `--spec-type draft-mtp` does not +fall back to plain generation. The server auto-selects a different speculative +implementation, which the MTP head model cannot satisfy: -Confirm from the response `timings`: if `draft_n` is `0`, MTP is not running. +``` +W srv load_model: [spec] failed to measure draft model memory: failed to create llama_context from model +I srv load_model: loading draft model '.../mtp-gemma-4-12B-it-Q4_0.gguf' +W common_speculative_init: draft model is specified but 'draft' speculative type is not explicitly enabled - enabling it +I common_speculative_impl_draft_simple: adding speculative implementation 'draft-simple' +I common_speculative_impl_draft_simple: - n_max=3, n_min=0, p_min=0.000000 +``` + +Startup then completes, and `GET /health` reports healthy. **Every completion +request fails:** + +``` +HTTP 500 +{"error":{"code":500,"message":"decode() failed: failed to process speculative batch","type":"server_error"}} +``` + +So the tell is not slow generation — it is a server that passes its health +check and cannot serve a single request. The earlier +`failed to create llama_context from model` line is logged at warning level and +is not treated as fatal. + +Reproduced with `data/repro-2026-08-07/mtp-g2.sh`; the response body above is +`data/repro-2026-08-07/G2.1.json`. -## Pascal caution 2 — draft backend sampling is non-repeatable on sm_61 +## Caution 2 — greedy output is not reproducible across requests by default -`--spec-draft-backend-sampling` defaults to **enabled** -(`common/common.h`: `backend_sampling = true`, "offload draft sampling to the -backend"). With it enabled on this P5200, repeated **greedy** requests with -identical input produced **different** outputs. +With `cache_prompt` at its default of `true`, repeated **identical greedy +requests** to the same server do not all return the same text. The first +request differs from the rest. -Passing `--no-spec-draft-backend-sampling` made repeated MTP runs byte-identical -to one another without materially reducing speed — the 56.44 tok/s above was -measured with backend sampling disabled. +The cause is prompt cache reuse changing the prompt batch split, not sampling: -This is reported as an observation on one sm_61 device, not as a diagnosis. It -has not been bisected, and no claim is made about other architectures. The -isolation runs are in `gemma12-mtp-cpu-sampler-1.json` and -`gemma12-mtp-cpu-sampler-2.json`. +| Request | `cache_n` | `prompt_n` | +| --- | ---: | ---: | +| 1st | 0 | 34 | +| 2nd and later | 7 | 27 | -Note that the flag is currently absent from the option list in -[`../speculative.md`](../speculative.md); it is added there by the same change -that adds this page. +A different batch shape gives a different floating-point reduction order, which +can flip a greedy token, after which the continuations diverge. + +Six conditions, five identical greedy requests each (`temperature: 0`, +`top_k: 1`, `seed: 42`): + +| | Configuration | Distinct outputs / 5 | +| --- | --- | ---: | +| A | MTP, draft backend sampling **enabled** (default) | 2 | +| B | MTP, draft backend sampling **disabled** | 2 | +| C | target only, no draft model (control) | 2 | +| D | target only, `-bs` (main backend sampling on) | 2 | +| E | target only, `cache_prompt: false` | **1** | +| F | MTP, `cache_prompt: false` | **1** | + +Three things worth reading off that table: + +- **A and B produced byte-identical output sets.** + `--no-spec-draft-backend-sampling` changed nothing. Backend sampling is not + the cause. +- **The control diverges too.** C has no draft model at all, so MTP is not the + cause either. C and D are likewise byte-identical to each other. +- **`cache_prompt: false` makes it fully repeatable**, in both the target-only + and the MTP condition. + +In every condition the split was the same: run 1 differed, runs 2–5 were +identical to one another — matching the cache-state table above exactly. + +If you need reproducible greedy output across requests, send +`"cache_prompt": false`, or ensure every request starts from the same cache +state. Nothing here is specific to sm_61; the same reasoning applies wherever +batch shape affects reduction order. ## What "deterministic" does and does not mean here -With `--no-spec-draft-backend-sampling`, all three MTP responses were -byte-identical **to one another**, and all three target-only responses were -byte-identical to one another. +Holding cache state fixed, MTP is repeatable: all five runs in condition F were +byte-identical. -The MTP response was **not** byte-identical to the target-only response. This is -expected rather than a defect: the target still verifies every proposed token, -but batched target evaluation can select a different greedy token than -single-token evaluation because GPU floating-point reduction order differs. +The MTP output was **not** byte-identical to the target-only output (condition F +vs condition E). This is expected rather than a defect: the target still +verifies every proposed token, but batched target evaluation can select a +different greedy token than single-token evaluation because GPU floating-point +reduction order differs. So this result demonstrates verified-token speculation and run-to-run repeatability. It should **not** be described as bit-identical to @@ -93,7 +143,6 @@ non-speculative generation. -md /path/to/mtp-gemma-4-12B-it-Q4_0.gguf \ -c 4096 -ngl all -ngld all -fa on --parallel 1 \ --spec-type draft-mtp --spec-draft-n-max 16 --spec-draft-p-min 0.9 \ - --no-spec-draft-backend-sampling \ --host 127.0.0.1 --port 8080 --no-webui --jinja ``` @@ -109,8 +158,8 @@ multi-slot measurement was taken. ## Raw data -[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/) — verbatim, as -produced by the run: +[`data/pascal-p5200-2026-08-02/`](data/pascal-p5200-2026-08-02/) — the +throughput run, verbatim: - `README.md` — the original run report, including a same-build `llama-bench` pass over Qwen3.6-35B-A3B Q4_K_M and Gemma 4 26B-A4B QAT Q4_0 on the same GPU @@ -121,3 +170,12 @@ produced by the run: - `gemma12-*-smoke.json`, `gemma12-mtp-cpu-sampler-*.json` — smoke and determinism-isolation runs - `qwen-*`, `gemma-*` — the `llama-bench` pass and its telemetry + +[`data/repro-2026-08-07/`](data/repro-2026-08-07/) — the caution repros: + +- `mtp-determinism.sh` — conditions A–D, including the no-draft control +- `mtp-determinism2.sh` — conditions E–F, the `cache_prompt: false` test +- `mtp-g2.sh` — caution 1, `-md` without `--spec-type` +- `{A..F}.{1..5}.json` — every response, with timings and cache counters +- `G2.1.json`, `G2.server.log.excerpt` — the 500 and the startup log +- `summary.txt` — condition-by-condition distinct-output counts