Repository navigation
Close M28 baseline and V2-0043 at iteration 26 - #2
Conversation
…rent performance vs FAST_V1 (V2-0002)
…frontier baseline (V2-0005)
- Implement zero-copy state handoff between Prefill V2 and Static Decode - Add PrefillV2Model::decode_step and generate for full autoregressive lifecycle - Route single-token matrix-vector projections to mx_repacked_mmv_kernel - Add bench/prefill_v2_e2e_bench verifying zero-copy continuity, multi-turn determinism, and P64/P512/P2048 + TG128 benchmarks - Record EXP-V2-0007 experiment and update architecture documentation
…HIP graph replay (EXP-V2-0008)
…x in decode (EXP-V2-0010)
… core and direct Q8 RMSNorm
…uning (EXP-V2-0012)
…aling in EXP-V2-0013
…e V2-0015 qualification
…am suffix attention
…ation benchmark (EXP-V2-0019)
…pport per-token streaming
…ulti-turn conversations
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It adds a substantial GPU memory-management and attention subsystem whose kernel implementations are not fully shown and raises a production output-token limit, warranting human review despite only minor doc-consistency issues being found.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
This PR is a checkpoint/closeout that preserves the immutable production candidate e68c0f20, closes the V2-0043 attention experiment as a PRIMARY GOAL PASS, and rejects the FP16 real-model KQ-fragment parity stretch variants. Functionally it wires a gated V2-0043 GQA attention path (with an optional real-input parity comparison against the tiled/GQA-tiled kernels) into the Qwen3.5 GPU pipeline, adds OpenAI chat stop-parameter parsing plus a raised output-token limit, extends the compatibility audit tool with recurrent/attention layer signature dumps, and reconciles current-state/experiment documentation. It fits the project's staged M0–M7 methodology by formalizing the M28 baseline before staging V2-0044.
Changes:
- Add environment-gated V2-0043 GQA attention path plus a debug real-input parity comparison harness in the GPU pipeline.
- Add OpenAI
stop(string/array) parsing and raise the max-output-token bound (4096→65536) with matching test updates. - Add/verify the
prefill_v2workspace and topology-block subsystems, extend the compatibility audit, and reconcile current-state docs and experiment records.
| File | Description |
|---|---|
| tools/qwen35_gpu_pipeline.hpp | Adds V2-0043 GQA attention path + real-input parity comparison against tiled/GQA-tiled kernels via env flags; shared split scratch reuse verified safe. |
| tools/m6a2_compatibility_audit.cpp | Dumps recurrent and attention layer tensor-type signatures grouped by layer for compatibility auditing. |
| tests/openai_api_test.cpp | Adds stop array/string parsing assertions and updates the invalid max_tokens boundary to the new 65537 limit. |
| src/prefill_v2/workspace.cpp | Contiguous arena allocation for prefill_v2 activation/attention buffers; sizing and pointer advances verified consistent. |
| src/prefill_v2/topology_block.cpp | Ping-pong GDN×3 + GQA block orchestration for forward/decode/profiled paths. |
| include/miinfer/prefill_v2/workspace.hpp | Workspace field contract; split-K comment still says 64 while allocation reserves 32 (flagged). |
| docs/current-state.md | Adds M28/V2-0043 checkpoint section; retained "Current experiment status — V2-0008" heading now contradicts it (flagged). |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| V2-0044 attention-frontier branch from that checkpoint. Preserve `e68c0f20`; | ||
| do not reopen the rejected V2-0043 numerical variants. | ||
|
|
||
| ## Current experiment status — V2-0008 (Dedicated Single-Token Decode Execution & Reusable HIP Graph Replay) |
| float* attn_k = nullptr; // [max_tokens, kKvDim (1024)] | ||
| float* attn_v = nullptr; // [max_tokens, kKvDim (1024)] | ||
| float* attn_gated_output = nullptr; // [max_tokens, kQDim (6144)] | ||
| float* splitk_attn_workspace = nullptr; // [max_tokens * 24 * 64 * (256 + 2)] for Split-K suffix attention |

Summary
e68c0f20and close V2-0043 PRIMARY GOAL PASS.Validation
The untracked
gpucore.3330886artifact is intentionally not part of this PR.