Repository navigation
bench(dpa): split-KV num_splits sweep + archived results (TLLM-ATTN-SPLITKV PR-D) - #14
Merged
Merged
Conversation
…weep (TLLM-ATTN-SPLITKV PR-D) Adds three split-KV variants per sweep value to the existing three-way harness: contiguous_splitkv / direct_splitkv / legacy_splitkv, invoked through the PR-B entry points. --num-splits accepts repeated or comma-separated values; 1 is a valid sweep point and directly measures the cost of the extra combine launch (design package section 2, open question 3). Correctness still precedes timing. The v1 bitwise gate (legacy vs direct vs contiguous) is unchanged; each split path additionally emits an equivalence_splitkv record against its same-kind single-pass output - bitwise at num_splits=1 (the regression anchor), max|diff| within the 2e-3 oracle tolerance above that. A failed anchor or tolerance check flips equiv_ok in shape_summary. One partial buffer sized for the sweep maximum is shared by all split paths (calls are serial on one stream). The convergence gate now covers every compared path; gather_k/gather_v stay diagnostic-only. JSONL schema bumps to tllm-dpa-kernel-bench-v2: provenance gains num_splits_sweep; sample/path_stats records carry num_splits; shape_summary gains splitkv_medians and splits equivalence into equiv_bitwise (v1 semantics) plus equiv_ok (extended gate). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…TKV PR-D)
32-shape x {1,2,4,8,16} num_splits sweep on RTX 5070 Ti, schema
tllm-dpa-kernel-bench-v2. Equivalence: 96/96 bitwise anchors at ns=1,
384/384 within 2e-3 tolerance at ns>1 (max_abs_diff 6.1e-05).
Key findings: crossover at visible ~256 for ns>=2; direct_splitkv@16 is
6.0x faster than single-pass direct at visible=2048 and makes direct
beat the full legacy path for the first time; up to ~1.9x regression at
visible<=129 from the fixed combine launch. ncu shows occupancy lifted
8.35% -> 13.08% (block-count ratio, not saturation) and SM throughput
0.86% -> 6.04%.
Decision: TLLM_ATTN_SPLITKV default stays off (single-pass) since the
small-window regression is real and the host cannot adapt per-length
without a device sync.
Only 4/32 shapes converged (CV<=10%), all at visible>=1024; reported
limits preserved. Raw JSONL and machine-readable summary included.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TLLM-ATTN-SPLITKV PR-D: extends
--dpa-benchwith split-KV paths and a--num-splits N[,M,...]sweep, then archives a full 32-shape benchmarkrun on RTX 5070 Ti (schema
tllm-dpa-kernel-bench-v2, v1 fields kept).contiguous_splitkv,direct_splitkv,legacy_splitkv) per num_splits value; equivalencegate runs before timing — ns=1 must be bitwise-identical to the
single-pass baseline (anchor), ns>1 uses a documented 2e-3 tolerance
for fp32 reduction-order differences.
equiv_bitwisekept for v1compat;
equiv_okcovers the whole matrix.raw JSONL + summary.json):
direct_splitkv@16is 6.0x faster than single-pass
directand for the first time makesdirect beat the full legacy path at kernel level.
combine launch (~5us).
saturation), SM throughput 0.86% -> 6.04%, identical DRAM traffic.
TLLM_ATTN_SPLITKVdefault stays off — the small-windowregression is real and the host cannot adapt per-length without a
device sync.
Limitations
shapes are sub-30us kernels on a noisy consumer GPU. Main conclusions
(crossover direction, long-context speedup) sit inside the converged
region; exact regression ratios at small S are directional only.
covered.
Validation
(ns>1), 0 failures — run before any timing.
./build/tiny_llm_tests: 225 passed.2eb97b2was clean (dirty_files=0in provenance);docs/results commit
c75b5edon top.Stack
Base: #13 (
tllm-attn-splitkv-wiring). Chain: #11 design -> #12 kernel-> #13 wiring -> this PR.
Test plan
tiny_llm_kernel_bench --dpa-bench --num-splits 1,2,4,8,16 --warmup 20 --reps 1000 --batch 100 --repeats 3 --clock-warmup 4 --out /tmp/x.jsonlreproduces schema v2 recordsGenerated with Devin