Summary
The full-attention phase uses g_activations as temporary storage for each block's online-softmax partial output. The indexing is per warp but not per block, so every attention block concurrently writes and reads the same global addresses.
This inter-block data race causes schedule-dependent hidden-state corruption and can change generated tokens. It affects the current main branch at commit bab084443f14e6ab3ba76f8e8e3f721ad2e56144.
Problem
During Phase 3 of full_attention_layer, each block writes its warp partials using:
g_activations[warp_id * FA_HEAD_DIM + lane_id * EPL + e] = out_acc[e];
__syncthreads();
Warp 0 then reads the same region to merge those partials:
fo[e] += g_activations[ww * FA_HEAD_DIM + lane_id * EPL + e] * s;
__syncthreads() only orders threads within one block. Because the address has no block or attention-head offset, blocks processing different query heads can overwrite the storage while another block is reading it.
Observed impact
The race was reproduced on an NVIDIA B200:
- A one-step run under warp-scheduling perturbation returned the same token but changed the hidden-state SHA-256, showing immediate internal corruption.
- At five decode steps, shared-memory delay and warp scrambling produced different token sequences from the baseline.
- A 30-step unperturbed baseline was nondeterministic in one of five repeated runs.
The reductions themselves use the same fixed warp tree for a fixed input. The divergence came from different blocks sharing this scratch region, rather than expected floating-point reduction-order variation.
Fix and validation
Using the existing per-block shared-memory scratch allocation as float storage for the attention merge removes the cross-block sharing. The allocation needs 4,096 floats:
16 warps * 256 values * sizeof(float) = 16 KiB
After replacing the two g_activations accesses with per-block shared-memory accesses and increasing MAX_ACT_DIM from 3,584 to 4,096:
- One-step baseline and warp-scrambler hidden-state hashes matched.
- Five-step baseline, shared-delay, and warp-scrambler token sequences matched.
- The 30-step baseline was deterministic across 8/8 runs.
- A DTester/NVBit confirmation found zero result mismatches across 574 kernel launches.
I have a minimal patch prepared for this change.
Research context
This issue and fix were identified, analyzed, and documented during a fine-grained, interactive debugging session with GitHub Copilot, powered by GPT-5.6. The work was conducted as part of a research project on GPU testing and verification that is currently under submission.
Summary
The full-attention phase uses
g_activationsas temporary storage for each block's online-softmax partial output. The indexing is per warp but not per block, so every attention block concurrently writes and reads the same global addresses.This inter-block data race causes schedule-dependent hidden-state corruption and can change generated tokens. It affects the current
mainbranch at commitbab084443f14e6ab3ba76f8e8e3f721ad2e56144.Problem
During Phase 3 of
full_attention_layer, each block writes its warp partials using:Warp 0 then reads the same region to merge those partials:
__syncthreads()only orders threads within one block. Because the address has no block or attention-head offset, blocks processing different query heads can overwrite the storage while another block is reading it.Observed impact
The race was reproduced on an NVIDIA B200:
The reductions themselves use the same fixed warp tree for a fixed input. The divergence came from different blocks sharing this scratch region, rather than expected floating-point reduction-order variation.
Fix and validation
Using the existing per-block shared-memory scratch allocation as
floatstorage for the attention merge removes the cross-block sharing. The allocation needs 4,096 floats:After replacing the two
g_activationsaccesses with per-block shared-memory accesses and increasingMAX_ACT_DIMfrom 3,584 to 4,096:I have a minimal patch prepared for this change.
Research context
This issue and fix were identified, analyzed, and documented during a fine-grained, interactive debugging session with GitHub Copilot, powered by GPT-5.6. The work was conducted as part of a research project on GPU testing and verification that is currently under submission.