Skip to content

Close M28 baseline and V2-0043 at iteration 26 - #2

Merged
GionaGranchelli merged 77 commits into
mainfrom
rewrite/m28-single-mi50-prefill
Sep 29, 2026
Merged

GionaGranchelli merged 77 commits into
mainfrom
rewrite/m28-single-mi50-prefill

Conversation

@GionaGranchelli

Copy link
Copy Markdown
Owner

Summary

  • Preserve immutable production candidate e68c0f20 and close V2-0043 PRIMARY GOAL PASS.
  • Record real-input FP16 KQ-fragment parity limitation and reject the stretch variants.
  • Reconcile current-state docs and persist clean-build, host/GPU test, short-generation, and exact P8192 smoke evidence.
  • Stage V2-0044 only after this checkpoint merges and post-merge checks pass.

Validation

  • Clean Release build passed (ROCm 7.2.1 / HIP Clang 20.0.0).
  • Host tests: 11/11 passed.
  • GPU tests: 9 passed; five model integration tests skipped because no model path was provided.
  • Fresh-build short generation passed.
  • Exact saved P8192 request processed 8,192 prompt tokens and one completion token in 38.141 s (smoke only; not a new performance qualification).

The untracked gpucore.3330886 artifact is intentionally not part of this PR.

- Implement zero-copy state handoff between Prefill V2 and Static Decode
- Add PrefillV2Model::decode_step and generate for full autoregressive lifecycle
- Route single-token matrix-vector projections to mx_repacked_mmv_kernel
- Add bench/prefill_v2_e2e_bench verifying zero-copy continuity, multi-turn determinism, and P64/P512/P2048 + TG128 benchmarks
- Record EXP-V2-0007 experiment and update architecture documentation
Copilot AI balanced review requested due to automatic review settings September 29, 2026 18:37
@GionaGranchelli
GionaGranchelli merged commit e0649c5 into main Sep 29, 2026
2 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It adds a substantial GPU memory-management and attention subsystem whose kernel implementations are not fully shown and raises a production output-token limit, warranting human review despite only minor doc-consistency issues being found.

Review effort: Balanced
Findings: 2 Low severity

Open (2)
What changed in this PR

This PR is a checkpoint/closeout that preserves the immutable production candidate e68c0f20, closes the V2-0043 attention experiment as a PRIMARY GOAL PASS, and rejects the FP16 real-model KQ-fragment parity stretch variants. Functionally it wires a gated V2-0043 GQA attention path (with an optional real-input parity comparison against the tiled/GQA-tiled kernels) into the Qwen3.5 GPU pipeline, adds OpenAI chat stop-parameter parsing plus a raised output-token limit, extends the compatibility audit tool with recurrent/attention layer signature dumps, and reconciles current-state/experiment documentation. It fits the project's staged M0–M7 methodology by formalizing the M28 baseline before staging V2-0044.

Changes:

  • Add environment-gated V2-0043 GQA attention path plus a debug real-input parity comparison harness in the GPU pipeline.
  • Add OpenAI stop (string/array) parsing and raise the max-output-token bound (4096→65536) with matching test updates.
  • Add/verify the prefill_v2 workspace and topology-block subsystems, extend the compatibility audit, and reconcile current-state docs and experiment records.
File Description
tools/​qwen35_gpu_pipeline.hpp Adds V2-0043 GQA attention path + real-input parity comparison against tiled/GQA-tiled kernels via env flags; shared split scratch reuse verified safe.
tools/​m6a2_compatibility_audit.cpp Dumps recurrent and attention layer tensor-type signatures grouped by layer for compatibility auditing.
tests/​openai_api_test.cpp Adds stop array/string parsing assertions and updates the invalid max_tokens boundary to the new 65537 limit.
src/​prefill_v2/​workspace.cpp Contiguous arena allocation for prefill_v2 activation/attention buffers; sizing and pointer advances verified consistent.
src/​prefill_v2/​topology_block.cpp Ping-pong GDN×3 + GQA block orchestration for forward/decode/profiled paths.
include/​miinfer/​prefill_v2/​workspace.hpp Workspace field contract; split-K comment still says 64 while allocation reserves 32 (flagged).
docs/​current-state.md Adds M28/V2-0043 checkpoint section; retained "Current experiment status — V2-0008" heading now contradicts it (flagged).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docs/current-state.md
V2-0044 attention-frontier branch from that checkpoint. Preserve `e68c0f20`;
do not reopen the rejected V2-0043 numerical variants.

## Current experiment status — V2-0008 (Dedicated Single-Token Decode Execution & Reusable HIP Graph Replay)
float* attn_k = nullptr; // [max_tokens, kKvDim (1024)]
float* attn_v = nullptr; // [max_tokens, kKvDim (1024)]
float* attn_gated_output = nullptr; // [max_tokens, kQDim (6144)]
float* splitk_attn_workspace = nullptr; // [max_tokens * 24 * 64 * (256 + 2)] for Split-K suffix attention
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants