Skip to content

[InlineSpiller] hoistSpillInsideBB: partial-subregister split shares spill slot→ full-width store corrupts loop-carried sibling (gfx950 miscompile) #225054

Description

@xgxanq

On gfx950 (CDNA3 / MI300-class), a fused multi-head-attention forward kernel (_attn_fwd, from ROCm aiter, compiled via Triton 3.8.0) is miscompiled by the register allocator. A
loop-carried buffer-descriptor value is silently overwritten by an unrelated value inside the K/V inner loop, so buffer_load reads K from the wrong address on every iteration
after the first.

The defect is in InlineSpiller: when greedy live-range-splits only a subset of the sub-registers of a wide value (sgpr_128), the resulting sibling shares the same spill stack
slot as the original, but a full-width store writes the undefined sub-registers back over the slot, clobbering the loop-carried value still live there.

This is the same class of hazard that #177703 (75f03a6) fixed for the cross-BB hoist path (HoistSpillHelper::isSpillCandBB). The symmetric same-BB path
(InlineSpiller::hoistSpillInsideBB) was not covered, and is the one this kernel hits.

Miscompile (observable)

Inside the K/V loop (.LBB0_12):

  • The K buffer descriptor s[0:3] is rebuilt at the loop top by v_readlane from VGPR lanes, feeding buffer_load_dwordx4 — source location mha.py:138 (descriptor definition).
  • The same lanes are later overwritten in-loop by v_writelane on the dropout / RETURN_SCORES path — source location mha.py:228 (Philox dropout store).
  • There is no in-loop save-back of those lanes before the back-edge, so the next iteration reads the dropout value as if it were the K descriptor address.

The .loc metadata is decisive: the lanes are defined by the descriptor (mha.py:138) but overwritten by dropout (mha.py:228, random.py), with no restore. The exact VGPR/lane is
not essential — the defect reproduces on different registers/lanes under different pressure layouts (it is register/lane-agnostic).

Root cause

llvm/lib/CodeGen/InlineSpiller.cpp shares one stack slot among all descendants of an Original value:

// Share a stack slot among all descendants of Original.
Original = VRM.getOriginal(edit.getReg());
StackSlot = VRM.getStackSlot(Original);

The K descriptor is an sgpr_128:

  • sub0_sub1 = buffer base address — loop-carried, spilled and reloaded every trip;
  • sub2_sub3 = buffer config constants.

The dropout path only needs the config, so greedy live-range-splits sub2_sub3 into a descendant vreg that only defines sub2_sub3 (sub0_sub1 is undef). Because the descendant
shares Original, the spiller reuses the same stack slot, then emits a full-width SI_SPILL_S128_SAVE, writing the undef sub0_sub1 over the loop-carried base address still living
in that slot. The next iteration reloads a corrupted descriptor.

The correctness precondition the spiller assumes — that descendants sharing a slot hold "equal values defined by Original's VNI" — is false for a partial-subregister split: the
descendant equals Original only on sub2_sub3, and differs (undef) on sub0_sub1, yet a full-width store touches the whole slot.

Three-layer evidence (Debug+asserts llc):

  1. post-greedy MIR: undef %N.sub0:sgpr_128 = ... immediately followed by SI_SPILL_S128_SAVE %N, %stack.K, where %stack.K is shared with the loop-carried descriptor's siblings.
  2. post-rewriter: only sub2_sub3 is redefined, but a full-width SI_SPILL_S128_SAVE of the whole sgpr_128 is emitted; sub0_sub1 has no nearby def → stale/undef.
  3. asm: dropout writes all four descriptor lanes; the next loop-top read gets the polluted lanes.

Relation to #177703

75f03a6 ("[InlineSpiller] Hoist spills only when all of its subranges are available") added, to HoistSpillHelper::isSpillCandBB (the cross-BB hoist path), a check that all
sub-ranges are live at the prospective slot index. The symmetric same-BB path, InlineSpiller::hoistSpillInsideBB, has no such guard. This kernel hits the same-BB path (spill
inside a self-looping loop body), so the #177703 fix does not apply.

Candidate fix

Mirror the #177703 guard on the same-BB path: in hoistSpillInsideBB, bail out of the hoist when not all sub-ranges of the sibling value are live at the hoist point, so the
full-width store that would write back undefined sub-registers is not emitted.

// All sub-ranges of the sibling value must be live at the hoist point; otherwise
// the full-width store would write back undefined sub-registers into the shared
// spill slot, corrupting the loop-carried sibling. Mirrors the isSpillCandBB fix (#177703).
bool partialDead = SrcLI.hasSubRanges() &&
!all_of(SrcLI.subranges(), [&](const LiveInterval::SubRange &SR) {
return SR.getVNInfoAt(Idx) != nullptr;
});
if (partialDead)
return false;

Measured effect (original kernel, controlled single-variable toggle). Same llc, the only variable is the guard on/off; input is the original _attn_fwd.llir:

┌───────┬────────────────────────────────────────────────────┐
│ Guard │ Final asm (dataflow judge) │
├───────┼────────────────────────────────────────────────────┤
│ off │ BUG DETECTED (K-desc lane clobbered, no save-back) │
├───────┼────────────────────────────────────────────────────┤
│ on │ CLEAN │
└───────┴────────────────────────────────────────────────────┘

The guard fires 97 times on this input, and the miscompile disappears.

Known gaps (disclosed up front)

  1. No reduced .ll test yet for this (hoist) path. The defect requires the allocator, under critical register pressure, to place the partial split's config and the non-restored
    dropout lane in the same slot with no in-loop rebuild — an emergent RA behavior. Hand-written minimal IR reaches the necessary conditions (shared-slot full-width store of a
    partial-def descendant) but not the sufficient one (the actual final-asm lane collision): such attempts are BUG at the MIR layer but CLEAN at the asm layer. llvm-reduce on the
    real kernel either drops the dropout store that creates the collision, or converges onto a different exit of the same root cause (see gap 2). Reproduction currently requires the
    original kernel.
  2. This root cause has a second exit not covered by the hoist guard. Under higher pressure (e.g. waves-per-eu raised, or llvm-reduce output), the polluting full-width store is
    emitted on the insertSpill / isRealSpill main path rather than via hoistSpillInsideBB. On such inputs the hoist guard still fires but does not change the final asm (measured: a
    reduced test triggers the guard 6 times yet stays BUG, versus 97 fires → CLEAN on the original kernel). A complete fix for that exit needs insertSpill to use a private slot for a
    partially-defined super-register vreg instead of the shared Original slot. Tracked separately.

Environment

  • Target: amdgcn-amd-amdhsa, -mcpu=gfx950, -O3.
  • Kernel specialization: IS_CAUSAL=1, BLOCK_M=256, BLOCK_N=64, BLOCK_DMODEL=64, RETURN_SCORES=1, ENABLE_DROPOUT=1, VARLEN=1 (dropout pulls in the Philox RNG, whose live state
    collides with the descriptor).
  • Bisect: the miscompile is exposed/hidden by 8a9c075 ([AMDGPU][CodeGen] Allow remat with multiple users in same region #214725), which is pure pressure perturbation, not a fix — it hides the bug on the original kernel but exposes it on a
    waves-per-eu-raised variant. The latent InlineSpiller defect is independent of that commit.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions