Skip to content

Full-attention scratch buffer races across blocks and corrupts decode output #3

Description

@tyler-utah

Summary

The full-attention phase uses g_activations as temporary storage for each block's online-softmax partial output. The indexing is per warp but not per block, so every attention block concurrently writes and reads the same global addresses.

This inter-block data race causes schedule-dependent hidden-state corruption and can change generated tokens. It affects the current main branch at commit bab084443f14e6ab3ba76f8e8e3f721ad2e56144.

Problem

During Phase 3 of full_attention_layer, each block writes its warp partials using:

g_activations[warp_id * FA_HEAD_DIM + lane_id * EPL + e] = out_acc[e];
__syncthreads();

Warp 0 then reads the same region to merge those partials:

fo[e] += g_activations[ww * FA_HEAD_DIM + lane_id * EPL + e] * s;

__syncthreads() only orders threads within one block. Because the address has no block or attention-head offset, blocks processing different query heads can overwrite the storage while another block is reading it.

Observed impact

The race was reproduced on an NVIDIA B200:

  • A one-step run under warp-scheduling perturbation returned the same token but changed the hidden-state SHA-256, showing immediate internal corruption.
  • At five decode steps, shared-memory delay and warp scrambling produced different token sequences from the baseline.
  • A 30-step unperturbed baseline was nondeterministic in one of five repeated runs.

The reductions themselves use the same fixed warp tree for a fixed input. The divergence came from different blocks sharing this scratch region, rather than expected floating-point reduction-order variation.

Fix and validation

Using the existing per-block shared-memory scratch allocation as float storage for the attention merge removes the cross-block sharing. The allocation needs 4,096 floats:

16 warps * 256 values * sizeof(float) = 16 KiB

After replacing the two g_activations accesses with per-block shared-memory accesses and increasing MAX_ACT_DIM from 3,584 to 4,096:

  • One-step baseline and warp-scrambler hidden-state hashes matched.
  • Five-step baseline, shared-delay, and warp-scrambler token sequences matched.
  • The 30-step baseline was deterministic across 8/8 runs.
  • A DTester/NVBit confirmation found zero result mismatches across 574 kernel launches.

I have a minimal patch prepared for this change.

Research context

This issue and fix were identified, analyzed, and documented during a fine-grained, interactive debugging session with GitHub Copilot, powered by GPT-5.6. The work was conducted as part of a research project on GPU testing and verification that is currently under submission.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions