diff --git a/.agents/coordination.md b/.agents/coordination.md index 842151c2..a2085715 100644 --- a/.agents/coordination.md +++ b/.agents/coordination.md @@ -132,6 +132,21 @@ leaves. Owns only NEW `.agents/specs/cli-chat-complete.md`, the fixture, or GPU/model-download change; verification is the CPU record/doc checker suite. The row and open-PR list were unclaimed at selection time. +**Scheduler recompute-preemption anchor-backfill spike (`ENG-PREEMPT-RECOMPUTE`, +2026-08-01, `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE`).** Codex (GPT-5), isolated +worktree `/home/mudler/.cache/codex-vllm-cpp-eng-preempt-spike`, branch +`codex/eng-preempt-recompute-spike`, base `upstream/main` `1448e981`. CPU-only +records/spec pass: inventory pinned-vLLM recompute-preemption dispatch, +request-state/KV lifecycle, matching upstream tests, exact local anchors, and a +row-sized follow-on breakdown for the already-shipped bounded implementation. +Owns only NEW `.agents/specs/preemption.md`, the `ENG-PREEMPT-RECOMPUTE` row, +this claim, the matching roadmap/status/benchmark current-state cells, +`.agents/parity-ledger.md`, and append-only `.agents/state.md`. No source, +header, build, test, fixture, model, kernel, README, lifecycle support claim, +or benchmark number changes. Verification uses the scheduler/request-queue CPU +tests when build tooling permits and all record/document checkers; no GPU, +model, download, compiler installation, or external execution host is needed. + **Canonical DONE-owner reachability repair (`KV-PREFIX-CACHE`, `SAMPLE-LOGPROBS`, `SPEC-DFLASH`, `MODEL-SPEC-qwen3-dflash-dflash-qwen3-for-causal-lm`, @@ -1328,6 +1343,7 @@ table, tests, CMake. Details in the state-log entry of the same date. | Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update | |---|---|---|---|---|---|---|---| | `CLAIM-SERVE-CLI-CHAT-SPIKE` | `SERVE-CLI-CHAT` (`INVENTORIED` -> `SPIKE`) | Codex (GPT-5) | isolated worktree `/home/mudler/.cache/sdd/localai-org-maint-bot-vllm.cpp/codex-serve-cli-chat-spike`; CPU-only records/spec, no GPU/model/download | branch `codex/serve-cli-chat-spike`, base `upstream/main` `1448e981` | Owns only NEW `.agents/specs/cli-chat-complete.md`; the `SERVE-CLI-CHAT` engine-matrix row and rollup; matching feature-matrix, porting-inventory, roadmap, STATUS, BENCHMARKS, coordination, ledger, and append-only state entries. No source/header/CMake/test/README/model/kernel/fixture edits. | `SPIKE` | 2026-08-01 - pinned-source inventory corrected: vLLM does ship remote `chat`/`complete`; accepted dual-mode design preserves the existing local invocation. CPU record/doc checker suite is the closing gate. | +| `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE` | `ENG-PREEMPT-RECOMPUTE` (`ANCHOR-BACKFILL` -> `SPIKE`) | Codex (GPT-5) | isolated worktree `/home/mudler/.cache/codex-vllm-cpp-eng-preempt-spike`; CPU-only records/spec, no GPU/model/download/compiler install | branch `codex/eng-preempt-recompute-spike`, base `upstream/main` `1448e981` | Owns NEW `.agents/specs/preemption.md`, the `ENG-PREEMPT-RECOMPUTE` row, this claim, roadmap/status/benchmark checkpoint text, parity-ledger and append-only state. No executable or README changes. | `SPIKE` | 2026-08-01 - pinned-source audit and W1-W4 contract written; agent-record and document checkers plus mutation suites pass. | | `CLAIM-LAGUNA-W1W2` | `MODEL-TEXT-laguna-laguna-for-causal-lm` (INVENTORIED→ACTIVE) | Claude Code (opus-4-8) | isolated worktree `.claude/worktrees/wf_2ae79b7a-246-1`; CPU-only `build-cpu` (`-DVLLM_CPP_CUDA=OFF` Release `-Werror`); NO GPU, NO 73 GB download — structural bring-up + oracle DECISION only | branch `spike/laguna-s21-w1w2`, base `main` `5c3da2f1` | Laguna-S-2.1 W1 oracle-decision + W2 structural bring-up. Owns ONLY: NEW `include/vllm/model_executor/models/laguna.h`, NEW `src/vllm/model_executor/models/{laguna_registry,laguna_weights,laguna}.cpp`, NEW `tests/vllm/models/test_laguna_scaffold.cpp`, NEW `.agents/specs/laguna-s21-w1w2-2026-07-30.md`, its two CMake registration lines (`CMakeLists.txt` source list + `tests/CMakeLists.txt`), the `LagunaForCausalLM` sorted-set + error-message insert in `tests/vllm/models/test_model_registry.cpp`, the `MODEL-TEXT-laguna-laguna-for-causal-lm` row (INVENTORIED→ACTIVE) + checklist rollup (ACTIVE 23→24 / INVENTORIED 285→284, engaged 42→43), and the record surfaces (this claim, roadmap breadth, docs/STATUS, docs/BENCHMARKS, parity-ledger, state). **NON-COLLISION:** additive TU + one REGISTER line ⇒ ZERO edit to any shared array; the forward is a `VT_CHECK(false)` W3 stub so no production path changes; MUST NOT touch README, Metal/SACRED/apex/darwin, or any other model/kernel source. | `DONE` | 2026-07-30 — **W1 oracle-decision + W2 structural bring-up LANDED (foreground, NOT pushed).** Registry (`laguna`/`LagunaForCausalLM`) + `ParseLagunaParams` (nested dual-rope + variable Q-head + ungrouped sigmoid-noaux MoE) + GGUF `blk.N.*` name-map + UD-Q4_K_XL quant-mix (ZERO new decode kernel) + KV-cache spec + per-layer forward-composition scaffold with reuse citations. `test_laguna_scaffold` 3/3·40 + `test_model_registry` 24/24; CPU full-library `-Werror` clean; record checkers rc=0. RESIDUALS (W3/W4): device materialization + real forward + the 3 new ops + strict dual-oracle gate on a fetched checkpoint. **SUPERSEDED by `CLAIM-LAGUNA-W3` (2026-07-31) which landed the real forward + the 3 new ops.** | | `CLAIM-LAGUNA-W3` | `MODEL-TEXT-laguna-laguna-for-causal-lm` (stays `ACTIVE`; W3 real forward + the 3 new ops landed; real-model dual-oracle gate still PENDING W4) | Claude Code (opus-4-8) | isolated worktree `.claude/worktrees/wf_43e61a78-0e2-1`; CPU-only `build-cpu` (`-DVLLM_CPP_CUDA=OFF` Release `-Werror`); NO GPU, NO 73 GB download — real forward CODE + unit gates only | branch `laguna-s21-w3`, base `main` `f2e463d5` (confirmed via `git rev-parse HEAD`) | Laguna-S-2.1 W3 — turn the W1/W2 `VT_CHECK(false)` forward stub into a REAL runnable host-reference composition + land the 3 genuinely-NEW small host ops. Owns ONLY: NEW `include/vllm/model_executor/models/laguna_ops.h` + `src/vllm/model_executor/models/laguna_ops.cpp` (softplus head-gate + ungrouped sigmoid-noaux router + dual per-layer RoPE cos/sin builders), the rewritten `src/vllm/model_executor/models/laguna.cpp` (`LagunaModel::Forward` real composition), the `LagunaParams` per-layer variable-Q-head helpers in `include/vllm/model_executor/models/laguna.h`, the new-op + forward unit cases appended to `tests/vllm/models/test_laguna_scaffold.cpp`, the `laguna_ops.cpp` line in `CMakeLists.txt`, NEW `.agents/specs/laguna-s21-w3-2026-07-31.md`, the `MODEL-TEXT-laguna-laguna-for-causal-lm` row cells + this claim, and docs/STATUS + docs/BENCHMARKS pointers. **NON-COLLISION:** file-disjoint from the concurrent MLA-fold lane (`mla_attention.cpp`/`deepseek_v2.cpp` untouched); additive `laguna_ops` TU + one CMake line; the loaders still `VT_CHECK(false)` so NO production/device path changes; MUST NOT touch README, Metal/SACRED/apex/darwin, or any other model/kernel source. | `ACTIVE` | 2026-07-31 — **W3 REAL forward + 3 new ops LANDED + UNIT-GATED (foreground, NOT pushed).** `laguna_ops.cpp`: `LagunaSoftplusHeadGate` (per-head softplus out-gate), `LagunaUngroupedRouterTopK` (sigmoid noaux_tc MINUS the group step + tie-break razor: lower index on equal choice, UNBIASED weights, renorm, routed_scaling), `BuildLaguna{FullYarn,Sliding}CosSin` (dual per-layer RoPE, reusing the pinned `compute_yarn_inv_freq` over the partial-64 dims). `LagunaModel::Forward` is now a REAL runnable f32 host-reference composition (variable-Q-head GQA + dual RoPE + sliding-window mask + softplus gate + dense L0 / ungrouped-MoE L1..47 + untied lm_head). `test_laguna_scaffold` **8/8·166** (softplus math; router selection + tie-break RED-first; dual-RoPE cos/sin bit-match vs hand ref; variable-Q-head shapes; forward composition on synthetic weights — RUNS, deterministic, gather==full-row, softplus gate wired) + `test_model_registry` 24/24; CPU full-library `-DVLLM_CPP_CUDA=OFF` `-Werror` clean; all record checkers rc=0. **HONEST residual (DEFERRED W4, needs the 73 GB checkpoint):** GGUF keep-quant tower materialization (loaders still LOUDLY throw) + device/paged production forward (runner variable-Q-head device wiring) + the strict dual-oracle greedy gate (llama.cpp-Q4_K token-exact + vLLM-NVFP4 near-tie). Risks: dual-RoPE numerics vs the fork on the real config; the reference forward is f32 whole-sequence (bf16 paged token-exactness is a W4 boundary); router tie-break vs the oracle's actual greedy selection. Row stays `ACTIVE`. | | `CLAIM-LAGUNA-W4` | `MODEL-TEXT-laguna-laguna-for-causal-lm` (stays `ACTIVE`; checkpoint fetched + fidelity corrected; real-model greedy gate is the W5 close) | Claude Code (opus-4-8) | isolated worktree `.claude/worktrees/wf_63140a7b-03e-1`; DGX `dgx.casa` GB10 for the 73.4 GiB fetch + GGUF metadata read + llama.cpp oracle build (foreground); CPU-verified fidelity corrections; NOT pushed | branch `worktree-wf_63140a7b-03e-1`, base `main` `570510a9` | Laguna-S-2.1 W4 — FETCH the UD-Q4_K_XL GGUF, read its metadata + tensor map AUTHORITATIVELY, and correct the fidelity errors the W1-W3 scaffold made from config.json guesses. Owns ONLY: `laguna.h`/`laguna_ops.{h,cpp}`/`laguna.cpp`/`laguna_weights.cpp` (QK-RMSNorm + `LagunaYarnMscale` + separate gate/up + verified name-map/quant-mix), the `test_laguna_scaffold.cpp` cases for those, NEW `.agents/specs/laguna-s21-w4-2026-07-31.md`, the `MODEL-TEXT-laguna-laguna-for-causal-lm` row cells + this claim, docs/STATUS + docs/BENCHMARKS pointers. **NON-COLLISION:** laguna-only additive edits; loaders still throw the keep-quant residual so NO production/device path changes; MUST NOT touch README, Metal/SACRED/apex/darwin, or any other model/kernel source. | `ACTIVE` | 2026-07-31 — **Checkpoint FETCHED + arch grounded in REAL bytes + 3 fidelity bugs fixed.** UD-Q4_K_XL GGUF (73.4 GiB, 3 shards, 814 tensors) fetched to dgx; metadata read authoritatively (arch `laguna`, `expert_gating_func=2` sigmoid, `expert_weights_scale=2.5`, `leading_dense_block_count=1`, rope factor 32/yarn_attn_factor 1.0, per-layer head_count `[48,72,72,72]`, quant mix: attn Q8_0 / experts gate-up Q4_K + down Q5_K / shared Q8_0 / router+norms F32). Fixed: (a) per-head QK-RMSNorm `attn_q/k_norm` (scope MISSED it — no config flag), (b) dual-RoPE mscale via llama.cpp `yarn_attn_factor·(1+0.1·ln(factor))` off GGUF factor 32 (not HF 128/1.4852), (c) SEPARATE `ffn_gate/up_exps`. Oracle = `poolsideai/llama.cpp@laguna` (mainline b10087+) same-quant; no vLLM-GGUF path for `laguna`. **HONEST residual (W5 close, needs the resident 73 GB run):** keep-quant tower materialization (`Mw`/`Sew` mirror of ds4) + host-orchestrated `ForwardGguf` (vt::MatmulBT/GemmRowSlice) + dual-RoPE inv_freq ramp bit-match + the real greedy run vs the llama.cpp same-quant oracle (token-exact or characterized near-tie). Row stays `ACTIVE`. | diff --git a/.agents/engine-matrix.md b/.agents/engine-matrix.md index 08ff9768..61c38dc9 100644 --- a/.agents/engine-matrix.md +++ b/.agents/engine-matrix.md @@ -36,7 +36,7 @@ forensics: roadmap_v1.md and the parity ledger. | Area | Rows | `ANCHOR-BACKFILL` | `PARTIAL` | `SPIKE` | `READY` | `ACTIVE` | `GATING` | `DONE` | `INVENTORIED` | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:| -| Engine and scheduling | 27 | 3 | 4 | 1 | 0 | 12 | 2 | 1 | 4 | +| Engine and scheduling | 27 | 2 | 4 | 2 | 0 | 12 | 2 | 1 | 4 | | KV cache and memory | 21 | 1 | 2 | 2 | 2 | 7 | 2 | 1 | 4 | | Parallelism | 6 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 5 | | Sampling and generation | 15 | 0 | 2 | 0 | 0 | 7 | 0 | 1 | 5 | @@ -46,7 +46,7 @@ forensics: roadmap_v1.md and the parity ledger. | LoRA and adapters | 2 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | | Long context and attention | 10 | 0 | 0 | 0 | 1 | 5 | 1 | 0 | 3 | | Loading, tokenizer, config | 9 | 1 | 3 | 0 | 0 | 2 | 1 | 1 | 1 | -| **Total** | **131** | **8** | **16** | **5** | **4** | **48** | **8** | **9** | **32** | +| **Total** | **131** | **7** | **16** | **6** | **4** | **48** | **8** | **9** | **32** | ## Engine core and scheduling @@ -56,7 +56,7 @@ forensics: roadmap_v1.md and the parity ledger. | `ENG-CHUNKED-PREFILL` | Basic token-budget chunked prefill | T0 | `vllm/config/scheduler.py:84`; `vllm/v1/core/sched/scheduler.py:835`; `tests/v1/core/test_scheduler.py:185,503,903` | `src/vllm/v1/core/sched/scheduler.cpp:225,548` | `tests/vllm/v1/test_scheduler.cpp:192`; `tests/vllm/models/test_qwen27_paged_forward.cpp:492` | `planned: specs/chunked-prefill.md` | `ANCHOR-BACKFILL` | - | | `KV-PREFIX-CACHE` | APC hashes, lookup, allocation, partial blocks, eviction, plus explicit/model-default cache policy. W0 ports arbitrary-group no-prefix coordination and makes hybrid/attention-free defaults cache-off. **Full-surface re-audit 2026-07-22 ([spike](specs/prefix-prompt-caching-parity.md)) — the ported core is DEEPER than this row read (chain hashing, pool, all three coordinators, the complete hybrid intersection and four single-type managers), and the residual gaps are narrower and DIFFERENT:** **`generate_block_hash_extra_keys`: W2 DONE 2026-07-27 (`CLAIM-ROADMAP-D4APC`)** — the hardcoded no-op is replaced by a 1:1 port of `kv_cache_utils.py:451-591` (`_gen_mm_extra_hash_keys` + LoRA name + `cache_salt`, fixed order lora->mm->salt; prompt_embeds deferred, no prompt-embeds path). `Request`/`EngineCoreRequest` carry `cache_salt` + `lora_name`; `FromEngineCoreRequest` sets them BEFORE the first hash (fixed a latent ordering bug: mm_features were assigned after the ctor already hashed). The latent correctness trap is CLOSED and RED-first proven: with the stub, a tenant-B request false-hits tenant-A's 48 cached tokens (`n1==48`); with extra keys `n1==0` (no false-share). This unblocks the MM + LoRA cache consumers. **prefix-cache statistics: CLOSED 2026-07-22** (W1) — `PrefixCacheStats`/`CachingMetrics` ported 1:1 with `log_stats` DEFAULTED ON, which unblocks the `BACKEND-GATE-CUDA-SGLANG-PREFIX` hit-proof requirement; first measured hit rate 0.75 on a repeated-prefix corpus; no `cache_salt`; 1 of upstream's 4 hash algos; `skip_reading_prefix_cache` absent; partial-block primitives throw (upstream's own are DEAD CODE — no caller in `vllm/` — so they are NOT owed as live behaviour). **Also cleared: the "blocked on a supported non-hybrid family" blocker is STALE** — dense models default APC ON and five have landed, yet NO gate has ever run cache-ON **MLA prefix-cache-hit assert fixed 2026-07-23** (`CLAIM-MLA-PREFIX-CACHE-ASSERT`): `FullAttentionManager::find_longest_cache_hit` asserted `kind()==kFullAttention`, aborting DeepSeek-V2 (MLA group, kind `kMlaAttention`, APC default-ON) under asserts-enabled builds — latent since `ec6f4be`, inert under Release/NDEBUG. Relaxed to upstream's precondition `isinstance(spec, FullAttentionSpec or ChunkedLocalAttentionSpec)` (single_type_kv_cache_manager.py:578-582; MLAAttentionSpec IS-A FullAttentionSpec) ⇒ accept `kFullAttention` / `kMlaAttention` / `kChunkedLocalAttention`; restores DeepSeek-V2 SACRED gate 8/8 asserts-on, full-attention byte-identical, new MLA prefix-cache-hit unit cases. | T0 | `vllm/config/model.py:1805-1860`; `vllm/engine/arg_utils.py:510,1160-1166,2473-2508`; `vllm/config/cache.py:39,93,95`; extra keys `vllm/v1/core/kv_cache_utils.py:539-574`; hasher factory `:673-730`; `vllm/v1/core/kv_cache_coordinator.py:377-425,782-834`; `tests/v1/core/test_prefix_caching.py:225,1475,2781` | hashes/managers `src/vllm/v1/core/kv_cache_utils.cpp:259,291`; **extra_keys** `generate_block_hash_extra_keys` + `_gen_mm_extra_hash_keys` `src/vllm/v1/core/kv_cache_utils.cpp`; `cache_salt`/`lora_name` on `include/vllm/v1/request.h` + `include/vllm/v1/engine/types.h`, copied in `src/vllm/v1/request.cpp` `FromEngineCoreRequest` (fields set before the first hash); `src/vllm/v1/core/kv_cache_manager.cpp:124`; no-prefix coordinator/factory `src/vllm/v1/core/kv_cache_coordinator.cpp:260,273,279,545`; model-default/hasher selection `src/vllm/entrypoints/model_loader.cpp:109,167,180,191`; CLI `examples/server/main.cpp:126`; **statistics** `include/vllm/v1/metrics/stats.h`, recorded `src/vllm/v1/core/kv_cache_manager.cpp:139-147`, reset flag `:270-276`, take-and-swap `make_prefix_cache_stats()`, per-step window fold at the end of `Scheduler::schedule()`, accessors `Scheduler`/`EngineCore`/`LLMEngine::prefix_cache_metrics()`; `Request::num_preemptions` un-deferred (`include/vllm/v1/request.h`, incremented in `Scheduler::preempt_request`) | existing APC primitives `tests/vllm/v1/test_kv_cache_utils.cpp:411,516,536`; no-prefix hybrid allocation/no-hit `tests/vllm/v1/test_kv_cache_coordinator.cpp:213`; default/override resolution `tests/vllm/entrypoints/test_loaded_engine_dense.cpp:343`; server help and online cache-off contracts `examples/CMakeLists.txt:34`; `tests/tools/test_online_gate_client.py:582,633`; statistics plus the first MEASURED hit rate `tests/vllm/v1/test_prefix_cache_stats.cpp` 12/12; **W2 extra_keys** — ported mm/lora/salt cases + ordering + hash-level no-false-share `tests/vllm/v1/test_kv_cache_utils.cpp` (29/29), manager-level salt-partition no-false-share (RED-proven `n1 48->0`) `tests/vllm/v1/test_kv_cache_manager.cpp` (10/10), CPU gate on dgx GB10. **W3 DONE 2026-07-27 (`CLAIM-ROADMAP-D4APC-W3`, dgx GB10, NOT pushed) — the FIRST-EVER cache-ON model gate:** `tests/parity/test_qwen3_apc_e2e.cpp` on `Qwen/Qwen3-4B` (dense, full-attention, APC-default-ON) 2/2 cases, 84/84 asserts — APC-ON hits 2240/2777 (rate 0.807) / APC-OFF 0; APC-ON == APC-OFF token-exact 5/6 (1 diff a vLLM-confirmed 0.125-nat near-tie); == vLLM-APC-ON teacher-forced (OFF 6/6 gap 0.0, ON 6/6 gap ≤0.125 nats, 0 outside top-20); TTFT 70.1→39.9 ms = 1.76×. NO engine code changed (gate-only over the already-shipped default-ON path); 4B SACRED 16/16 no-regression. Oracle vLLM 0.25.0. Ledger: [parity-ledger.md#L746](parity-ledger.md#L746) | [prefix-prompt-caching-parity.md](specs/prefix-prompt-caching-parity.md) (umbrella); [prefix-caching.md](specs/prefix-caching.md) (cache-policy leaf) | `DONE` (dense APC path; W4 events/W5 partial/W6 mamba-align/W7 reset endpoint tracked in `KV-EVENTS`/`KV-MAMBA-ALIGN`/own future rows) | `a41af480` | | `KV-PREFIX-MATCH-UNIT` | `--prefix-match-unit` (config `prefix_match_unit`): the finest token boundary a prefix-cache hit can land on == the `hash_block_size`/"prefix match unit" the block hasher uses. NEW in 0.26 (absent at the prior `e24d1b24`/0.25.0 pin). For a HYBRID/multi-group model the resolver `resolve_kv_cache_block_sizes` computes `hash_block_size = prefix_match_unit if set else gcd(group_block_sizes)` (scheduler block size = `lcm`), letting matching land FINER than a physical block (e.g. 16/32 tokens inside a 1024-token block) provided every group block size is divisible by it; single-group (dense) models ignore the knob. Backs off to the scheduler block size when no prefix-cache/connector consumer is active or a mamba group diverges from `cache_block_size` (mamba_cache_mode != "align"); throws on a non-divisible unit. **W0 spike + W1 resolver LANDED 2026-07-28 (`CLAIM-PREFIX-MATCH-UNIT`, NOT pushed):** `resolve_kv_cache_block_sizes` ported 1:1 (explicit-parameter signature vs upstream's `VllmConfig`, our config surface is threaded), RED-first unit-gated (default gcd `!=` `=16` override). `PARTIAL`: the config/CLI/ABI field (W2), the scheduler threading of a resolved `hash_block_size != block_size` + mamba partial-tail stop (W3, needs the `KV-BLOCK-POOL` align path that still throws), and the benchmark (W4) are deferred. Default path byte-identical (single-group inert; scheduler still passes `block_size`). | T1 | `vllm/engine/arg_utils.py:696,1222,1940`; `vllm/config/cache.py:56-67`; resolver `vllm/v1/core/kv_cache_utils.py:626-688`; hasher `:691-748`; call site `vllm/v1/engine/core.py:154`; scheduler `vllm/v1/core/sched/scheduler.py:76,268-270,282,312-318`; fine-grained view `vllm/v1/core/single_type_kv_cache_manager.py:683,697` | resolver `src/vllm/v1/core/kv_cache_utils.cpp:638` (`resolve_kv_cache_block_sizes`), decl `include/vllm/v1/core/kv_cache_utils.h`; hash_block_size already plumbed `get_request_block_hasher` `src/vllm/v1/core/kv_cache_utils.cpp:577`; DEFERRED align path throws `src/vllm/v1/core/block_pool.cpp:93,220` (shared with `KV-BLOCK-POOL`) | `tests/vllm/v1/test_prefix_match_unit.cpp:64,88,99,119,129,145,164,186` 8/8 (29 assertions): single-group inert + DCP scale, multi-group default=gcd, `=16` override finer-than-default (RED), finer-than-1024-block, non-divisible throws, no-consumer back-off + connector-alone re-enable, mamba non-align back-off vs align gcd, hasher-granularity RED (coarse 2 vs fine 4 hashes); [parity-ledger.md](parity-ledger.md) | [prefix-match-unit.md](specs/prefix-match-unit.md) | `PARTIAL` | `CLAIM-PREFIX-MATCH-UNIT` | -| `ENG-PREEMPT-RECOMPUTE` | FCFS tail preemption with recompute | T0 | `vllm/v1/core/sched/scheduler.py:1142`; `tests/v1/core/test_scheduler.py:930` | `src/vllm/v1/core/sched/scheduler.cpp:102,157`; `src/vllm/v1/core/sched/request_queue.cpp:36` | `tests/vllm/v1/test_scheduler.cpp:247,295`; `tests/vllm/v1/test_request_queue.cpp:91` | `planned: specs/preemption.md` | `ANCHOR-BACKFILL` | - | +| `ENG-PREEMPT-RECOMPUTE` | FCFS/priority KV-pressure preemption with recompute. The bounded local path frees KV, resets computed progress, prepends the victim, reports events/reset ids, and resumes it as an MRV2 new request; spike audit found stale speculative-token cleanup, prefix-cache on/off recomputation coverage, and encoder-cache/in-flight cleanup still open | T0 | `vllm/v1/core/sched/scheduler.py:427-443,563-613,1203-1245`; `tests/v1/core/test_scheduler.py:1016-1072` @ `555967922` | `src/vllm/v1/core/sched/scheduler.cpp:129-148,328-383,562-586`; `src/vllm/v1/core/sched/request_queue.cpp:30-49` | `tests/vllm/v1/test_scheduler.cpp:370-408,420-518,895-959,1091-1155`; `tests/vllm/v1/test_request_queue.cpp:113-129` | [preemption.md](specs/preemption.md) | `SPIKE` | `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE` | | `ENG-CUDAGRAPH` | Decode graph capture/replay modes (host-cluster cleanup: capture-size set derived from `max_num_seqs` mirroring vLLM `_set_cudagraph_sizes`; 2026-07-18 graph-baked-scratch use-after-free fix — the 35B c2+ online-serving IMA blocker) | T0 | `vllm/config/compilation.py:53,1319,683-684,1438-1444`; `vllm/config/vllm.py:1667-1770`; `vllm/v1/worker/gpu/cudagraph_utils.py:116`; `tests/compile/test_config.py:122,229` | `src/vt/cuda/cuda_backend.cu:76,97,105`; `include/vllm/model_executor/models/decode_graph_sizes.h`; `src/vllm/model_executor/models/qwen3_5.cpp:3754,3952`; `src/vllm/v1/worker/gpu/runner.cpp:577,597`; graph-safe scratch (retire-on-grow so graph-baked scratch pointers stay valid) `src/vt/cuda/graph_safe_scratch.h`, `src/vt/cuda/cuda_moe_marlin.cu:75`, `src/vt/cuda/cuda_matmul_nvfp4.cu:766`, `src/vt/cuda/cuda_matmul_nvfp4_cutlass.cu:105`, `src/vt/cuda/cuda_matmul_fp8_cutlass.cu:95` | `tests/vt/test_cuda_backend.cpp:98`; `tests/vllm/models/test_decode_graph_sizes.cpp`; `tests/vt/test_graph_safe_scratch.cpp`; explicit 35B gate `tests/parity/test_qwen36_paged_engine.cpp:140` | [blocktable-host-cluster-cleanup.md](specs/blocktable-host-cluster-cleanup.md); [decode-graph-scratch-uaf-2026-07-18.md](specs/decode-graph-scratch-uaf-2026-07-18.md) | `PARTIAL` | - | | `ENG-BATCH-INVARIANT` | Opt-in deterministic execution across scheduler batch sizes (`VLLM_BATCH_INVARIANT=1`): batch-invariant matmul/norm/attention/collectives plus persistent-scheduler NVFP4; production default remains off | T1 | default/env `vllm/envs.py:89,576-578`; initialization `vllm/v1/worker/gpu_worker.py:1262`; NVFP4 dispatch `csrc/libtorch_stable/quantization/fp4/nvfp4_scaled_mm_sm120_kernels.cu:212-220`; suite fixture `tests/v1/determinism/conftest.py:9-12`; operator/e2e `tests/v1/determinism/test_nvfp4_batch_invariant_scaled_mm.py`, `tests/v1/determinism/test_nvfp4_batch_invariant.py` @ `702f481` | - | [W3-C3R executed contract](specs/nvfp4-persistent-plan-cache.md#w3-c3r-batch-shape-localization-and-gate-correction-2026-07-13): production-default ours and vLLM both change outputs across batch shapes; no local opt-in implementation is claimed | `planned: specs/batch-invariant-execution.md` | `INVENTORIED` | - | | `ENG-ASYNC-SCHED` | Async/overlap scheduling (AsyncScheduler placeholders + depth-2 batch-queue step + async D2H on a copy stream); vLLM's DEFAULT at the pin — mirror obligation per B3. **Host-side machinery + runner device-input half + sampler-OUTPUT half LANDED + CPU-gated (2026-07-16):** `AsyncScheduler` placeholder accounting, `step_with_batch_queue` depth-2, `ResolveAsyncScheduling` default-ON-when-compatible + `MaxConcurrentBatches`, `VT_ASYNC_SCHED` rollback; the runner device-input path `combine_sampled_and_draft_tokens`; PLUS the sampler-OUTPUT half — `vt::Backend` event/pinned primitives (`AllocPinned`/events, CUDA cudaHostAlloc+cudaEvent, CPU sync-degeneration), `AsyncGPUModelRunnerOutput` (device sampled-id snapshot → non-blocking D2H on a copy queue + event; `get_output()` waits only that event; MAIN queue never blocked), `Sampler::forward(sampled_ids_out)` device-resident greedy, `GPUModelRunner::sample_tokens_async` + `runner_supports_async`, and the `Executor`+`step_with_batch_queue` seam resolving `get_output()` at CONSUME time. All behind `VT_ASYNC_RUNNER`/`set_async_input_combine`, default OFF. Sync path byte-identical (placeholder sites INERT while count 0; combine off; `sample_tokens_async` degenerates to sync when async off; `sampled_ids_out=nullptr`). **ENABLE-FLIP LANDED + CPU-gated (2026-07-16):** (1) `LoadedEngine` now reorders `runner_` before the scheduler and builds an `AsyncScheduler` + `max_concurrent_batches=2` when `ResolveAsyncScheduling(runner_.runner_supports_async())` resolves ON (else the byte-identical synchronous `Scheduler` + depth-1); the resolved mcb threads into `AsyncLLM`→`EngineCoreProc` (`step_with_batch_queue`) and the "Asynchronous scheduling is enabled/disabled" log mirrors vLLM for A/B audit; (2) the device combine/scatter kernel (`_combine_sampled_and_draft_tokens_kernel` + last_sampled scatter) is ported to CUDA (`src/vt/cuda/cuda_combine_tokens.cu`), main-stream-ordered on the CUDA async path so it DELETES `sample_tokens_async`'s pre-scatter `Synchronize`; the CPU backend keeps the host loop. `VT_ASYNC_RUNNER=1` engages full W3; `VT_ASYNC_SCHED=0` is the same-binary rollback. Production default (no env) stays synchronous byte-identical. **FULL W3 DGX proof RAN twice** — `f086b64` (5/5 gates PASS; c16 TPOT −5.4 ms WIN, tput neutral, TTFT +36 % = Little's-law repayment) and the 2026-07-16 re-proof on the THROUGHPUT-lever fix (persistent pooled sampled-id/pinned buffers + `Sampler` greedy scratch removing ALL per-step `cudaMalloc`/`cudaFree`/`cudaHostAlloc`/event-create from the sampled-id path, incl. the overlap-killing `cudaFree` inside `get_output`; mirrors `gpu_model_runner.py:873-878` + `async_utils.py:12-70`): token-exactness **6/6 PASS**, interleaved c16 **tput −0.32 % (gate ≥+1.5 % FAILS), TPOT −4.95 ms retained, TTFT +34.8 %** — the allocator lever is REFUTED as the tput unlock (≤0.1 % of a ~165 ms c16 step). **DEFAULT FLIPPED ON 2026-07-17** (`VT_ASYNC_RUNNER` default ON via the pure `AsyncRunnerFlagIsOn` predicate, mirroring `vllm/config/vllm.py:992-1044`): the discriminator (`6ea7856`) proved vLLM's own async pays the identical +26–31 % TTFT / −0.7 to −0.9 % tput / −2.6 to −4.3 ms TPOT envelope and W3-ON nets positive (both binding ITL-tail anomalies flip to PASS), so the "needs a throughput lever" ship-gate is RETIRED — W3 is a parity/mirror obligation with a tails+TPOT win. The flip is TOKEN-NEUTRAL (async-ON ≡ async-OFF bit-identical on DGX). `VT_ASYNC_RUNNER=0` = runner-level rollback, `VT_ASYNC_SCHED=0` = scheduler-level rollback. TTFT means rise into vLLM's async envelope BY DESIGN — the next binding grid runs async by default and its TTFT must NOT be misread as a regression. **ROBUSTNESS FIX 2026-07-20 (`discard_request_mask`):** the runner was missing vLLM's `discard_request_mask`, so `GPUModelRunner` emitted a sampled token for prefill-CHUNK requests too; under async this drained a `num_output_placeholders` never reserved (the `is_prefill_chunk` path adds none) → the `async_scheduler.cpp` `num_output_placeholders >= 0` assertion aborted on c8 + short-output (chunked prefill + preemption). FIX mirrors vLLM: `execute_model` computes `exec_state_.discard[i] = seq_len < num_tokens` (`gpu_model_runner.py:2048`); `sample_tokens` clears those rows to empty (`outputs.py:303`), the async path passes `invalid_req_indices` to `AsyncGPUModelRunnerOutput::get_output` (`gpu_model_runner.py:3625` + `outputs.py:303`). Scheduler UNCHANGED (assertion kept — it was correct once the runner honors `scheduler.py:1888-1890`). Sync/non-chunked decode byte-identical (mask all-zero); DGX 27B 235/235 + 35B 315/315, `vllm-bench` c8+short-output+chunked+kv-pressure no longer crashes, memcheck 0. Ledger [parity-ledger.md](parity-ledger.md) 2026-07-20 row | T1 | `vllm/v1/core/sched/async_scheduler.py:12`; `vllm/config/vllm.py:490,990,1038`; `vllm/v1/engine/core.py:519`; `vllm/v1/worker/gpu/input_batch.py:304-406`; `vllm/v1/worker/gpu/async_utils.py:12-70`; `vllm/v1/worker/gpu/gpu_model_runner.py:242-332`; `vllm/v1/outputs.py:298-307` | `src/vllm/v1/core/sched/async_scheduler.cpp:10,45`; placeholder plumbing `src/vllm/v1/core/sched/scheduler.cpp:148,164,605`; `src/vllm/v1/engine/core.cpp:91` (`step_with_batch_queue`, async-output seam); `src/vllm/v1/engine/core_proc.cpp:32,46`; config `include/vllm/config/scheduler.h:117,165,188`, `src/vllm/config/scheduler.cpp:12`; `include/vllm/v1/request.h:187`; runner input leaf `src/vllm/v1/worker/gpu/prepare_inputs.cpp`, `src/vllm/v1/worker/gpu/input_batch.cpp`; runner output leaf `include/vt/backend.h`+`src/vt/backend.cpp`+`src/vt/cuda/cuda_backend.cu` (event/pinned), `include/vllm/v1/worker/gpu/async_output.{h,cpp}` (`AsyncGPUModelRunnerOutput`), `src/vllm/v1/sample/sampler.cpp` (`sampled_ids_out`), `src/vllm/v1/worker/gpu/runner.cpp` (`sample_tokens_async`/`runner_supports_async`), `src/vllm/v1/executor/executor.cpp`+`include/vllm/v1/worker/gpu/model_runner_base.h` (async seam); enable-flip `include/vllm/entrypoints/model_loader.h`+`src/vllm/entrypoints/model_loader.cpp` (`runner_` before scheduler, `ResolveAsyncEnabled`/`MakeScheduler`, `AsyncScheduler`+mcb=2, log), `include/vllm/v1/engine/async_llm.h`+`src/vllm/v1/engine/async_llm.cpp` (mcb param → `EngineCoreProc`); device kernel `include/vt/cuda/combine_tokens.h`+`src/vt/cuda/cuda_combine_tokens.cu`, wired `src/vllm/v1/worker/gpu/runner.cpp` (CUDA combine/scatter branch removes the pre-sync) | `tests/vllm/v1/test_async_scheduler.cpp:1` (6 cases, 54 asserts; RED vs base Scheduler 2/6 fail); depth-2 engine cycle `tests/vllm/v1/test_engine_core_proc.cpp:479` (mcb=2, async-output seam); config resolution `tests/vllm/test_scheduler_config.cpp:75`; enable-flip construction matrix `tests/vllm/entrypoints/test_loaded_engine_dense.cpp` (runner×VT_ASYNC_SCHED → scheduler type + mcb; RED = un-flipped engine, 3/3 ON-arm asserts fail); runner input leaf `test_combine_tokens.cpp` (RED = stale → 5/7 fail), `test_input_batch.cpp`, `test_runner.cpp` (async-ON≡sync); output leaf `tests/vt/test_backend.cpp` (event/pinned contract), `tests/vllm/v1/worker/test_async_output.cpp` (materialize/flush/snapshot; RED = +1 splice), `test_runner.cpp` (`sample_tokens_async` decode ≡ sync); full CPU ctest 111/111, tools 164/164. Prior diagnostic `3812d8` six-leg control: total **1.002153×**, TTFT **0.862159×**, no GPU-time reduction (neutral for speed). **DEFAULT-FLIP (2026-07-17):** new pure CPU flag test [test_async_runner_flag.cpp](../tests/vllm/v1/worker/test_async_runner_flag.cpp) (11 asserts, default-ON/'0'-off); construction matrix [test_loaded_engine_dense.cpp](../tests/vllm/entrypoints/test_loaded_engine_dense.cpp) INVERTED (default → AsyncScheduler+mcb=2; RED verified 5 asserts fail vs un-flipped). CPU clean `-Werror` rebuild, full serial ctest **116/116**, tools **164/164**. **DGX re-confirmation** (evidence `dgx:~/work/vllm.cpp-async-flip`, CUTLASS+FA2 hard-verified, one flock): shipping default (async ON + RMSNorm-fast OFF) → **27B 235/235 + 35B 315/315** with the "Asynchronous scheduling is enabled (mcb=2)" log, and both rollback arms (`VT_ASYNC_RUNNER=0`, `VT_ASYNC_SCHED=0`) 235/235 + 315/315 log "disabled"; async arms BIT-IDENTICAL (token-neutral). Closing record [parity-ledger.md#L502](parity-ledger.md#L502) | [async-serving.md](specs/async-serving.md) | `DONE` | `6ea7856` | diff --git a/.agents/parity-ledger.md b/.agents/parity-ledger.md index 1156d192..16e9b123 100644 --- a/.agents/parity-ledger.md +++ b/.agents/parity-ledger.md @@ -886,3 +886,4 @@ Columns: | 2026-07-31 (`SERVE-C-ABI` W0 contract spike; `CLAIM-SERVE-C-ABI-SPIKE`; CPU-only records/docs) | Accepted `.agents/specs/c-api-library.md` for the already-shipped original C packaging layer: complete scope, vLLM semantic chain/deviation, ABI v10/19-symbol baseline, ownership/error/version/dispatch rules, exact code/test anchors, gates, dependencies, risks, and W1-W5 follow-ons. Also fixes the verified stale public `VLLM_ABI_VERSION 9` labels in README/USAGE to the source-of-truth v10 and adds the missing v10 usage-table entry. No production/test/CMake source changed. | Pinned vLLM `555967922` has no C ABI; behavior beneath the adapter remains owned by its vLLM-derived engine rows. The flat ABI is the recorded llama.cpp-style packaging deviation and may translate, never reimplement, policy. | **CPU/records gate only; benchmark NOT APPLICABLE.** Focused C11/C++/dlopen/export gate passed 3/3 after explicitly building `vllm_shared`; five record checkers pass. `check-agent-record` reports the base tree's same six missing closing-commit objects (`444ea9d7`, `7a3f04b2`, `164453a2`), none in this row/diff. Row stays `ANCHOR-BACKFILL` because all-symbol dlsym coverage (chat symbols currently omitted), historical-layout compatibility, allocation-failure no-throw proof, lifetime sanitizer stress, and a standalone real-model C consumer remain W1-W5. | | 2026-07-31 (`CLAIM-CPU-GCC12-WERROR-PORTABILITY`; maintenance, rows `QUANT-GGUF-KEEPQ-LOADER` + `KV-OFFLOAD`; lifecycle unchanged) | Removes two GCC 12 production-library `-Werror` blockers without suppressions: the GGUF prefault keeps the same one-byte-per-page volatile XOR but uses simple assignment, and the KV filesystem tier builds the identical `...tmp` suffix with append operations inside its thread-local initializer. No API, algorithm, default, CUDA, fixture, or golden change. | Behavior remains grounded in the accepted loader and KV-persistence leaf specs: llama.cpp mmap prefault intent and vLLM `tiering/fs/io.py` unique temporary-file publication. This is compiler portability, not a parity-surface change. | RED: GCC 12 failed first at `qwen3_5_gguf_weights.cpp:49` (`-Wvolatile`), then at `fs_io.cpp:66` (`-Wrestrict`). GREEN: production `vllm` and focused test targets build clean; focused CTest 2/2 (`test_gguf_keep_quant`, `test_kv_offload_fs`). Full all-target build is PARTIAL at 42% on unrelated test-only GCC 12 `-Wrestrict` diagnostics in `test_deepseek_v2_paged_engine.cpp` and `test_glm4_moe_lite_paged_engine.cpp`; no full-CTest claim. Benchmark NOT APPLICABLE. | | 2026-08-01 (`SERVE-CLI-CHAT` W0 contract spike; `CLAIM-SERVE-CLI-CHAT-SPIKE`; CPU-only records/spec) | Accepts `.agents/specs/cli-chat-complete.md`, corrects the inventory from “no direct commands” to the actual pinned `chat`/`complete` surface, and decomposes a dual-mode port: exact remote OpenAI HTTP/SSE commands plus preservation of the existing in-process invocation as a compatibility alias. No production, test, CMake, model, kernel, fixture, or generated file changes. | Pinned vLLM `5559679229`: command registration `vllm/entrypoints/cli/main.py:17-37,73-98`; model/auth resolution and stream shaping `vllm/entrypoints/cli/openai.py:30-100`; chat `:155-234`; complete `:237-312`. The local compatibility baseline is `examples/cli/main.cpp:1-207`. | CPU record/doc gates only; benchmark `NOT APPLICABLE`, `benchmark_binding=false`. Implementation remains absent and the row moves `INVENTORIED` -> `SPIKE`. W1-W5 name parse, transport, complete, chat, and packaging gates, including fake-server request/SSE transcript parity, Release `-Werror`, ASan+UBSan, and TSan. | +| 2026-08-01 (`ENG-PREEMPT-RECOMPUTE` W0 spike; `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE`; CPU-only records) | Accepts `.agents/specs/preemption.md` for the already-shipped bounded recompute-preemption path and moves its record from `ANCHOR-BACKFILL` to `SPIKE` without changing runtime behavior or support. The audit confirms KV release, state reset, FCFS/priority victim selection, front retry, event/reset-id reporting, and MRV2 resumed-as-new output, while naming stale speculative-token cleanup, prefix-cache recomputation tests, and encoder cleanup as open. | Pinned vLLM `scheduler.py:427-443,563-613,1203-1245` and `tests/v1/core/test_scheduler.py:1016-1072` @ `555967922`; local anchors and W1-W4 map are in the accepted spike. | Records-only, `benchmark_binding=false`. Existing CPU test anchors cover the bounded path; no new source/test claim is made. Record/document checkers are the binding W0 gate. W1/W2 are CPU-verifiable; encoder cleanup is dependency-gated and the final both-model every-axis gate requires a future uncontended GPU host. | diff --git a/.agents/roadmap_v1.md b/.agents/roadmap_v1.md index ee8d437e..4d1a8383 100644 --- a/.agents/roadmap_v1.md +++ b/.agents/roadmap_v1.md @@ -489,6 +489,12 @@ are ALREADY multimodal architectures brought up text-only, so the track COMPLETE models we already ship + benchmark. Full seam map + M0–M5 W-plan: [multimodal-track.md](specs/multimodal-track.md). +`ROAD-V1-A` scheduler record checkpoint (2026-08-01): +`ENG-PREEMPT-RECOMPUTE` moves `ANCHOR-BACKFILL` to `SPIKE` under the accepted +[preemption contract](specs/preemption.md). No runtime/support claim advances; +W1/W2 are CPU leaves, encoder cleanup waits on its inventoried dependencies, +and the binding both-model GPU gate remains W4. + | Order | Block | Big area / outcome | Canonical detailed table | Spike coverage | State | Next gate | |---:|---|---|---|---|---|---| | 0 | `ROAD-V1-A` | Restore exact performance closure against the faster applicable vLLM v0.25.0/SGLang floor before broader roadmap implementation | [`BACKEND-GATE-CUDA-VLLM`](backend-matrix.md), [`BACKEND-GATE-CUDA-SGLANG`](backend-matrix.md), [`BACKEND-GATE-CUDA-SGLANG-PREFIX`](backend-matrix.md), [`SERVE-GATE-ONLINE`](engine-matrix.md), [`KV-PREFIX-CACHE`](engine-matrix.md), [`KV-MAMBA-ALIGN`](engine-matrix.md), [`KV-DEVICE-RESIDENCY`](engine-matrix.md), [`SERVE-ASYNC-LLM`](engine-matrix.md), [`KERNEL-GEMM-BF16`](kernel-matrix.md), [`KERNEL-GEMM-NVFP4-W4A4`](kernel-matrix.md), [`KERNEL-ATTN-FA2`](kernel-matrix.md), [`KERNEL-GDN-PACKED-DECODE`](kernel-matrix.md), [`KERNEL-GDN-AOT-BF16`](kernel-matrix.md), [`SERVE-STREAM-USAGE`](engine-matrix.md), [`SERVE-E2E-NIGHTLY`](engine-matrix.md), [benchmark protocol](benchmark-protocol.md) | v0.25.0 target `702f481` is audited. **NEW BINDING `9ecd9d0`: 114/124** (gate NO — 10 remain; full production default set = async + Triton GDN cubin + bit-identical fast RMSNorm + bit-identical fast gated-RMSNorm; supersedes `a875397`'s 52/124, `246a23c`'s 49/124, `3f256ab`'s 55/124). Mem 4/4, c1 20/20, c2 20/20, c16 19/20, c4 & c32 18/20, c8 15/20. The bit-identical (0-ulp) fast decode-kernel stack (`348d12d`+`9ecd9d0`) closed +62 axes — confirming the decode deficit was norm/quant/act kernel glue. Landed: async (`a0013a2`, `ENG-ASYNC-SCHED` DONE), vendored Triton GDN cubin (`a321d7c`), packed-decode equivalence (`e47b4d6`), qkvz (`45f9e6d`), windowed-load memory PASS. SGLang remains open; the SGLang behavior-parity lane (`ENG-SGLANG-BEHAVIOR-FLAG` cache-aware LPM admission + `KV-SGLANG-RADIX-CACHE` radix alias) is IMPLEMENTED CPU-side 2026-07-27 (`CLAIM-SGLANG-IMPL`, rows `ACTIVE`) — the `--schedule-policy=lpm` admission the `BACKEND-GATE-CUDA-SGLANG-PREFIX` gate exercises now exists, output-neutral | `PARTIAL` | **EFFECTIVE PARITY-OR-BETTER reached (27B vs vLLM)** (two-grid totality 115/124: 110 pass-in-both + 5 coin-flip; the 9 residuals are the low-conc-median edge of a net-positive determinism tradeoff — we win the tails + high-conc + throughput). No closeable real deficit. NEXT: accept the 27B verdict → 35B performance closure → SGLang floor → open the T1/T2 portfolio (orders 1–14). **35B ENGINE LEVER LANDED 2026-07-19 (`ENG-MOE-SHARED-AUX`, `CLAIM-MOE-SHARED-AUX-1`):** the first slice of the multi-stream intra-step OVERLAP (the largest c1/c2 lever) — MoE shared-expert MLP on an aux CUDA stream concurrent with the routed experts, byte-identical, default-ON; in-situ 35B TPOT A/B WINS every concurrency (c1 −5.6% … c32 −1.5%, zero regression). Orchestrator re-grids the binding c1/c2/c4; remaining overlap slices + portable glue-fusion continue the 35B closure | diff --git a/.agents/specs/preemption.md b/.agents/specs/preemption.md new file mode 100644 index 00000000..435ae66e --- /dev/null +++ b/.agents/specs/preemption.md @@ -0,0 +1,168 @@ +# ENG-PREEMPT-RECOMPUTE spike: scheduler recompute preemption + +Date: 2026-08-01 +Pin: vLLM `5559679229bc961848b121ccdeaa8fa5d79bec98` +Row: `ENG-PREEMPT-RECOMPUTE` +Claim: `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE` + +## Scope + +This spike inventories the running-request recompute-preemption path. The +bounded baseline is FCFS tail eviction when KV allocation fails: release the +victim's KV blocks, mark it `PREEMPTED`, reset computed-token progress, prepend +it to the waiting queue, report its id, and later resend it as a new request. + +In scope for row closure are FCFS and priority victim selection, undoing work +already scheduled in the current step, speculative-token cleanup, encoder +cache/in-flight-prefill cleanup, output/event bookkeeping, recomputation with +prefix caching on and off, and resume through the MRV2 runner contract. +Disaggregated preemption, swap preemption, pipeline-parallel stale outputs, and +forced administrative reset are separate mechanisms and are out of scope. + +The current PR changes records only. It does not advance a support claim. + +## Upstream chain + +Pinned vLLM owns this behavior entirely in the host scheduler: + +- `vllm/v1/core/sched/scheduler.py:427-443,471-520` derives per-request work + from `num_tokens_with_spec - num_computed_tokens` and schedules running + requests first. +- `vllm/v1/core/sched/scheduler.py:563-613` retries KV allocation, selects the + FCFS tail or lowest-priority victim, restores same-step token/encoder budgets, + and stops only after the current request evicts itself. +- `vllm/v1/core/sched/scheduler.py:1203-1225` frees KV and encoder state, + removes the request from in-flight prefills, clears speculative tokens, + resets computed progress, increments metrics, prepends the victim, and emits + the reset id. +- `vllm/v1/core/sched/scheduler.py:1227-1245` advances computed progress only + after a schedule is formed, which is why resetting to zero means recompute. +- `vllm/v1/core/request_queue.py:131-199` defines FCFS prepend/pop ordering. + +No runtime-selected dependency kernel participates. A later performance gate +must still trace both engines because recomputation changes the subsequent GPU +workload, but there is no kernel dispatch decision to resolve in this spike. + +## Our baseline + +The bounded FCFS path is already present: + +- `src/vllm/v1/core/sched/scheduler.cpp:129-148` frees KV, changes status, + resets progress, increments/logs preemption, prepends the request, and records + the reset id. +- `src/vllm/v1/core/sched/scheduler.cpp:328-383` retries allocation and selects + FCFS or priority victims, including same-step token-budget rollback. +- `src/vllm/v1/core/sched/scheduler.cpp:419-520` resumes waiting/preempted + requests, and `:562-586` folds resumed requests into MRV2 new-request output. +- `src/vllm/v1/core/sched/request_queue.cpp:30-49` implements FCFS prepend and + removal. + +Existing CPU evidence covers KV exhaustion, FCFS-front retry, event/count +bookkeeping, and resumed-as-new behavior. Four upstream obligations remain +open, so the row cannot be `DONE`: + +1. `preempt_request` does not clear `request->spec_token_ids`. +2. Encoder-cache and `_inflight_prefills` cleanup have no equivalent state in + this bounded scheduler path. +3. The local same-step rollback does not restore encoder compute allocations; + the encoder scheduling surface is not yet ported here. +4. No focused test proves prefix-cache-enabled recomputation or preservation of + an already-sampled output token across preemption, both exercised upstream. + +The first gap is locally implementable and must be RED-tested. Gaps 2 and 3 +depend on the encoder/cross-attention scheduler rows. Gap 4 is CPU-testable. + +## Port map + +| Upstream | Local target | Disposition | +|---|---|---| +| `scheduler.py:563-613` | `src/vllm/v1/core/sched/scheduler.cpp:328-383` | Present for FCFS/priority and token-budget rollback; encoder rollback deferred | +| `scheduler.py:1203-1225` | `src/vllm/v1/core/sched/scheduler.cpp:129-148` | KV/status/progress/metrics/queue present; spec-token and encoder cleanup incomplete | +| `scheduler.py:1227-1245` | `src/vllm/v1/core/sched/scheduler.cpp:562-586` | Present through the MRV2 resumed-as-new fold | +| `request_queue.py:131-199` | `src/vllm/v1/core/sched/request_queue.cpp:30-49` | Present | + +The only deliberate structural deviation is C++ ownership: `Scheduler` owns +requests in a map and queues hold borrowed pointers. Behavior and ordering +remain vLLM-defined. + +## Tests to port + +| Upstream test | Local tier | Current disposition | +|---|---|---| +| `tests/v1/core/test_scheduler.py:1016-1072::test_preempt_during_execution` | CPU doctest | Partial: `tests/vllm/v1/test_scheduler.cpp:370-408` covers eviction/reset, but not the upstream sampled-output preservation tail | +| `tests/v1/core/test_scheduler.py` priority preemption/resumption cases | CPU doctest | Present at `tests/vllm/v1/test_scheduler.cpp:895-959,1091-1155`; keep as cross-policy coverage | +| FCFS prepend behavior used by `_preempt_request` | CPU doctest | Present at `tests/vllm/v1/test_request_queue.cpp:113-129` | +| Spec-token clearing on preemption | CPU doctest | Missing; add a RED-first case before the implementation change | +| Prefix-cache-enabled recomputation | CPU doctest | Missing; port with caching on/off parametrization and exact recomputed-token/block assertions | +| Encoder-cache/in-flight-prefill cleanup | CPU doctest | SKIPPED until encoder scheduling state exists; tracked by the dependent encoder rows | + +## Gates + +W1 and W2 are fully CPU-verifiable: + +```sh +cmake -S . -B build-cpu -G Ninja -DCMAKE_BUILD_TYPE=Release \ + -DVLLM_CPP_CUDA=OFF -DVLLM_CPP_SERVER=OFF +cmake --build build-cpu --target test_scheduler test_request_queue -j2 +ctest --test-dir build-cpu --output-on-failure \ + -R '^(test_scheduler|test_request_queue)$' +``` + +Run the record gates at every checkpoint: + +```sh +python3 scripts/check-agent-record.py +python3 tests/scripts/test_agent_record.py +python3 scripts/check-doc-checkpoint.py --staged +python3 tests/scripts/test_doc_checkpoint.py +python3 scripts/check-readme-structure.py +python3 scripts/check-model-checklist.py +python3 scripts/check-fusion-consistency.py +python3 scripts/check-device-leakage.py +python3 scripts/check-env-doc.py +``` + +Correctness closure additionally requires an identical seeded request stream +against pinned vLLM with caching on and off, comparing schedule outputs, +statuses, recomputed-token counts, reset ids, and final token ids. The eventual +performance/memory closure uses both gate models at the standard large- +concurrency workload, preemption forced by an identical KV budget, with nsys on +both engines and every-axis comparison under the benchmark protocol. Those GPU +gates keep the row below `DONE`; they are not required for this records-only +spike. + +## Dependencies + +- W1/W2 depend only on `ENG-SCHED-CORE` and `KV-MANAGER-ALLOC`, whose bounded + implementations already exist; no model, GPU, network data, or new license is + required. +- W3 depends on the encoder/cross-attention scheduler surface (`KV-CROSS-ENCODER-SPECS` + and `ATTN-ENCODER-CROSS`) and cannot be silently folded into this row early. +- W4 requires the pinned vLLM oracle, both gate models, and an uncontended GPU + host selected by developer preferences. + +## Work breakdown + +| Leaf | Scope | Files | Gate | +|---|---|---|---| +| W1 | Complete the upstream basic case: preserve sampled output, clear stale spec tokens, assert reset id/event/count | scheduler implementation plus `test_scheduler.cpp` | CPU focused tests | +| W2 | Add prefix-cache on/off recomputation and resume-totality cases | `test_scheduler.cpp`, optionally KV-manager test helpers | CPU focused tests + pinned scheduler differential | +| W3 | Add encoder-cache/in-flight-prefill cleanup and same-step encoder-budget restoration after its dependencies land | scheduler/encoder manager and ported upstream tests | CPU oracle fixtures, then multimodal e2e | +| W4 | Close token, memory, latency, and throughput parity under forced preemption | parity/e2e harness and benchmark records only | both gate models, vLLM, nsys, every-axis grid | + +W1 and W2 are sequential because they share the scheduler test file. W3 is +dependency-blocked. W4 follows semantic closure. + +## Risks and decisions + +- Recompute means resetting computed progress, not deleting sampled output. + Tests must keep these two token domains distinct. +- Prefix caching may reduce the amount physically recomputed after reset; this + is expected upstream behavior, not a reason to demand a full-prompt miss. +- A preempted request can have been scheduled earlier in the same step. Budget + restoration is part of correctness, not an optimization. +- We considered closing only the already-tested FCFS subset, splitting every + policy into separate rows, or keeping one upstream-semantic row. The selected + design keeps one row and uses W leaves because FCFS and priority share the + same `_preempt_request` state transition; separate lifecycle claims would + hide cross-policy cleanup gaps. diff --git a/.agents/state.md b/.agents/state.md index 173f48b6..80f1cc66 100644 --- a/.agents/state.md +++ b/.agents/state.md @@ -34505,3 +34505,19 @@ The required Slack selection notification was attempted through the bundled secret-safe sender to the only conventional target available, `#general`, but Slack returned `channel_not_found`. No channel ID/name is configured and no credential was inspected or exposed. + +## 2026-08-01 - ENG-PREEMPT-RECOMPUTE W0 spike + +`CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE`, isolated worktree +`/home/mudler/.cache/codex-vllm-cpp-eng-preempt-spike`, branch +`codex/eng-preempt-recompute-spike`, base `upstream/main` `1448e981`. +Accepted `.agents/specs/preemption.md` and moved the existing bounded scheduler +row from `ANCHOR-BACKFILL` to `SPIKE`, without changing executable files or a +support claim. The source audit confirmed FCFS/priority KV-pressure victim +selection, KV free, computed-progress reset, front retry, events/reset ids, and +MRV2 resumed-as-new output. It also found concrete closure gaps: stale +`spec_token_ids` are not cleared, prefix-cache on/off recomputation and sampled- +output preservation lack focused parity cases, and encoder-cache/in-flight +cleanup waits on the encoder scheduler surface. Resume with W1: port the +missing upstream test tail, add a RED spec-token case, then implement the +one-line cleanup and run the focused CPU scheduler/request-queue gates. diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index fc665a30..e780aaaa 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -42,6 +42,18 @@ Two byte-exact fronts on the GB10 NVFP4 decode graph (`~/laguna-xs-nvfp4`, ids ` The 78 removed casts are ~78 µs/step — BELOW the ~0.3 ms nsys run-to-run noise (the dominant `enable_if` projection GEMV drifts ±0.4–0.7% between runs), so GPU-busy reads at parity; the win is the deterministic node-count drop (fewer graph nodes → less ramp/drain the CUDA graph doesn't hide), reflected as a neutral-to-+0.29% `decode_hp`. Lands default-ON on the same "shrink the captured node count, byte-exact" basis as the glue/preamble/addnorm folds (`=0` is a same-binary A/B opt-out). The remaining decode tail (projection GEMV ~55%, lm_head 7.4%, we-win Marlin MoE ~12%) is at/beyond parity; the load-time Marlin repack (`TransposeToInt32`/`gptq_marlin_repack`/`ProcessScales`, 20046 inst) is a one-time cost, correctly zeroed by the 2-length diff (the nsys-aggregate trap). +## Scheduler recompute-preemption W0 spike (2026-08-01) - NOT APPLICABLE + +`ENG-PREEMPT-RECOMPUTE` / `CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE` is a +records-only checkpoint. It inventories the already-shipped bounded scheduler +path and its missing parity gates without changing source, tests, build rules, +runtime behavior, workload, or accepted support state. No throughput, latency, +or memory number is applicable (`benchmark_binding=false`). The next +reproduction is the CPU W1 focused scheduler/request-queue gate in +[the spike](../.agents/specs/preemption.md); a binding benchmark remains W4 and +requires forced preemption on both gate models against pinned vLLM on an +uncontended GPU host. + ## GCC 12 production-library portability (2026-07-31, `CLAIM-CPU-GCC12-WERROR-PORTABILITY`) - NOT APPLICABLE / all-target build PARTIAL This maintenance checkpoint changes no runtime algorithm or benchmark axis. diff --git a/docs/STATUS.md b/docs/STATUS.md index b8bfbbca..eb075fca 100644 --- a/docs/STATUS.md +++ b/docs/STATUS.md @@ -31,6 +31,14 @@ citing "vLLM 0.25.0" are the last binding measurement against the prior oracle ## Capability status +Scheduler recompute preemption (2026-08-01, +`CLAIM-ENG-PREEMPT-RECOMPUTE-SPIKE`) is `SPIKE`, with no support-state change. +The existing CPU path covers KV-pressure victim selection, KV release, reset, +front-of-queue retry, metrics, and MRV2 resume. The accepted +[spike](../.agents/specs/preemption.md) identifies the current closure gaps: +stale speculative-token cleanup, prefix-cache on/off recomputation coverage, +and encoder-cache/in-flight cleanup after the encoder scheduler surface lands. + GCC 12 production-library maintenance (2026-07-31): the two known `-Werror` blockers in GGUF prefaulting and KV-offload temporary-file naming are fixed without behavior or lifecycle changes. The production library and focused