Skip to content

bench(dpa): split-KV num_splits sweep + archived results (TLLM-ATTN-SPLITKV PR-D) - #14

Merged
holtwood merged 2 commits into
masterfrom
tllm-attn-splitkv-bench
Sep 15, 2026
Merged

holtwood merged 2 commits into
masterfrom
tllm-attn-splitkv-bench

Conversation

@holtwood

Copy link
Copy Markdown
Member

Summary

TLLM-ATTN-SPLITKV PR-D: extends --dpa-bench with split-KV paths and a
--num-splits N[,M,...] sweep, then archives a full 32-shape benchmark
run on RTX 5070 Ti (schema tllm-dpa-kernel-bench-v2, v1 fields kept).

  • Harness: 3 split-KV fetch variants (contiguous_splitkv,
    direct_splitkv, legacy_splitkv) per num_splits value; equivalence
    gate runs before timing — ns=1 must be bitwise-identical to the
    single-pass baseline (anchor), ns>1 uses a documented 2e-3 tolerance
    for fp32 reduction-order differences. equiv_bitwise kept for v1
    compat; equiv_ok covers the whole matrix.
  • Results (docs/performance/results/2026-09-15-rtx5070ti-splitkv.md +
    raw JSONL + summary.json):
    • 480/480 equivalence records pass; max_abs_diff 6.1e-05.
    • Crossover at visible ~256 (ns>=2); at visible=2048 direct_splitkv@16
      is 6.0x faster than single-pass direct and for the first time makes
      direct beat the full legacy path at kernel level.
    • Small windows regress (up to ~1.9x at visible<=129) due to the fixed
      combine launch (~5us).
    • ncu: occupancy 8.35% -> 13.08% (matches 14->112 blocks / 70 SMs, not
      saturation), SM throughput 0.86% -> 6.04%, identical DRAM traffic.
  • Decision: TLLM_ATTN_SPLITKV default stays off — the small-window
    regression is real and the host cannot adapt per-length without a
    device sync.

Limitations

  • Only 4/32 shapes converged (CV<=10%), all at visible>=1024; the small
    shapes are sub-30us kernels on a noisy consumer GPU. Main conclusions
    (crossover direction, long-context speedup) sit inside the converged
    region; exact regression ratios at small S are directional only.
  • Kernel-level medians only — no TTFT/TPOT/serving claims.
  • num_splits=32, other block sizes, and multi-KV-head geometries not
    covered.

Validation

  • Equivalence gate: 96 bitwise anchors (ns=1) + 384 tolerance records
    (ns>1), 0 failures — run before any timing.
  • ./build/tiny_llm_tests: 225 passed.
  • compute-sanitizer initcheck smoke on the bench paths: 0 errors.
  • clang-format clean (local v19; CI runs 18.1.8).
  • Benchmark commit 2eb97b2 was clean (dirty_files=0 in provenance);
    docs/results commit c75b5ed on top.

Stack

Base: #13 (tllm-attn-splitkv-wiring). Chain: #11 design -> #12 kernel
-> #13 wiring -> this PR.

Test plan

  • CI green on this branch
  • Re-run: tiny_llm_kernel_bench --dpa-bench --num-splits 1,2,4,8,16 --warmup 20 --reps 1000 --batch 100 --repeats 3 --clock-warmup 4 --out /tmp/x.jsonl reproduces schema v2 records

Generated with Devin

holtwood and others added 2 commits September 15, 2026 12:02
…weep (TLLM-ATTN-SPLITKV PR-D)

Adds three split-KV variants per sweep value to the existing three-way
harness: contiguous_splitkv / direct_splitkv / legacy_splitkv, invoked
through the PR-B entry points. --num-splits accepts repeated or
comma-separated values; 1 is a valid sweep point and directly measures
the cost of the extra combine launch (design package section 2, open
question 3).

Correctness still precedes timing. The v1 bitwise gate
(legacy vs direct vs contiguous) is unchanged; each split path
additionally emits an equivalence_splitkv record against its same-kind
single-pass output - bitwise at num_splits=1 (the regression anchor),
max|diff| within the 2e-3 oracle tolerance above that. A failed anchor
or tolerance check flips equiv_ok in shape_summary.

One partial buffer sized for the sweep maximum is shared by all split
paths (calls are serial on one stream). The convergence gate now covers
every compared path; gather_k/gather_v stay diagnostic-only.

JSONL schema bumps to tllm-dpa-kernel-bench-v2: provenance gains
num_splits_sweep; sample/path_stats records carry num_splits;
shape_summary gains splitkv_medians and splits equivalence into
equiv_bitwise (v1 semantics) plus equiv_ok (extended gate).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…TKV PR-D)

32-shape x {1,2,4,8,16} num_splits sweep on RTX 5070 Ti, schema
tllm-dpa-kernel-bench-v2. Equivalence: 96/96 bitwise anchors at ns=1,
384/384 within 2e-3 tolerance at ns>1 (max_abs_diff 6.1e-05).

Key findings: crossover at visible ~256 for ns>=2; direct_splitkv@16 is
6.0x faster than single-pass direct at visible=2048 and makes direct
beat the full legacy path for the first time; up to ~1.9x regression at
visible<=129 from the fixed combine launch. ncu shows occupancy lifted
8.35% -> 13.08% (block-count ratio, not saturation) and SM throughput
0.86% -> 6.04%.

Decision: TLLM_ATTN_SPLITKV default stays off (single-pass) since the
small-window regression is real and the host cannot adapt per-length
without a device sync.

Only 4/32 shapes converged (CV<=10%), all at visible>=1024; reported
limits preserved. Raw JSONL and machine-readable summary included.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@holtwood
holtwood changed the base branch from tllm-attn-splitkv-wiring to master September 15, 2026 04:25
@holtwood
holtwood merged commit 4d1977f into master Sep 15, 2026
2 checks passed
@holtwood
holtwood deleted the tllm-attn-splitkv-bench branch September 16, 2026 06:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant