Repository navigation
[AIROCMLIR-1292] [EXTERNAL] Cherry-pick the upstream fix for the float-based 32-bit div/rem expansion - #511
Conversation
071cbf2 to
b0ffc28
Compare
umangyadav
left a comment
There was a problem hiding this comment.
Check first if following LLVM PRs fix the issue or not. If they do then cherry pick them
llvm/llvm-project#201186
llvm/llvm-project#202753
b0ffc28 to
ea746a5
Compare
Yeah, those are the fixes, I cherry-picked both onto this PR and dropped my workaround |
For each cherry pick we should have patch and patch context. |
ea746a5 to
b4e0d1a
Compare
Sorry, I didn't know those had to be separate. Split it the way you described so one [EXTERNAL] commit per cherry-pick carrying only the upstream code, and the patch files plus their context in a commit of their own |
There was a problem hiding this comment.
Verdict: COMMENT · Findings: 1 (0 Critical, 0 Major, 1 Minor)
Scope
Cherry-picks two upstream AMDGPU backend fixes into the vendored Triton-pinned LLVM tree (external/llvm-project/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp plus 13 affected LLVM CodeGen lit expectations), and records them as llvm-patches/patch201186.patch / patch202753.patch with matching entries in llvm-patches/llvm-patch-content.txt.
The pair narrows the float-based div/rem expansion (renamed expandDivRem24 -> expandDivRemToFloat) from 24 bits to 23 signed / 22 unsigned, because trunc(float(Num) * rcp(float(Den))) only self-corrects an underestimated quotient and the 1 ulp error of v_rcp_f32 can overestimate well below 0x800000. This fixes the silent wrong results and intermittent memory-access faults in makeGroupedGridLayout's bid % thisMBlocksPerGroup once the grid exceeds ~8.3M workgroups, without touching GridLayoutEmitter.cpp.
Findings
Verification performed, all clean:
- Both vendored-tree commits carry the required
[EXTERNAL]subject prefix; the third commit touches onlyllvm-patches/and correctly does not. - Each
.patchrecord matches its corresponding[EXTERNAL]commit diff (modulo theexternal/llvm-project/path prefix and hunk offsets from the pinned tree's context), and both have index entries with upstream commit, PR, files, symptom, root cause, and a "drop on the bump that includes ..." line. - Patch names sort so the bump guide's
sort -rreverse-applies202753before201186, which the record documents as required. - Final vendored state is verbatim upstream; the rename is complete and no stale
expandDivRem24reference remains. - rocMLIR back-port check: every touched path (
external/llvm-project/,llvm-patches/) is on the rocmlirTriton-only list, so no back-port note is required.
One Minor suggestion on regression coverage at llvm-patches/llvm-patch-content.txt:801.
Notes
- Fidelity nit, no action wanted: upstream's own comment and assert message in
expandDivRemToFloatImplsay0x40000where the declaration comment correctly says[-0x400000,0x3FFFFF], and read "must be <= than". Keeping the cherry-pick byte-identical to upstream is the right call here; worth an upstream follow-up rather than a local divergence. - Worth watching: narrowing the unsigned bound 24 -> 22 pushes more runtime-divisor sites off the cheap float path onto the full 32-bit expansion (clearly visible in the
udiv.i32.llexpectation churn). The PR description notes the group-local division stays on the float path via its 16-bit mask, and Jenkins is green, so this looks contained — but it is a broad backend change affecting every kernel's index math, not just the GEMM grid layout. - The bonus coverage of
makeGxNGridLayout'sbid / (splitKV * mBlocks)in the attention path is a genuine argument for fixing this in the backend rather than in the grid-layout emitter.
CI status
No failing or cancelled checks. Jenkins, Build and Test, and MIGraphX all pass; Code coverage is still pending and review is this pipeline's own in-progress check.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The SelectionDAG fallback still permits the unsafe 24-bit float expansion when CodeGenPrepare is bypassed.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Cherry-picks LLVM AMDGPU fixes that restrict unsafe float-based integer div/rem expansion.
Changes:
- Limits float expansion to signed 23-bit and unsigned 22-bit operands.
- Updates AMDGPU code-generation tests and adds boundary coverage.
- Records upstream patch provenance.
| File | Description |
|---|---|
llvm-patches/llvm-patch-content.txt |
Documents both upstream fixes. |
external/llvm-project/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp |
Tightens div/rem expansion bounds. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/amdgpu-codegenprepare-idiv.ll |
Adds boundary tests. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/sdiv.ll |
Updates signed division checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/sdivrem24.ll |
Updates signed div/rem checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/udiv.ll |
Updates unsigned division checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/udivrem24.ll |
Updates unsigned div/rem checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/urem64.ll |
Updates remainder lowering checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/GlobalISel/udiv.i32.ll |
Updates i32 GlobalISel checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/GlobalISel/udiv.i64.ll |
Updates i64 GlobalISel checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/GlobalISel/urem.i32.ll |
Updates i32 remainder checks. |
external/llvm-project/llvm/test/CodeGen/AMDGPU/GlobalISel/urem.i64.ll |
Updates i64 remainder checks. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
b4e0d1a to
4e22f15
Compare
4e22f15 to
092e014
Compare
…m cherry-picks Co-authored-by: Cursor <cursoragent@cursor.com>
092e014 to
05fc5d7
Compare

Motivation
GEMM kernels silently returned wrong results, and intermittently died with a memory access fault, once the grid grew past roughly 8.3 million workgroups. Found while bisecting a gfx90a weekly tuning failure: a 9216x1536 f16 GEMM with a 1x1 tile produced 238592 wrong output elements on every run and faulted on roughly one run in twelve. 55 of the 6466 problem/tile pairs in the tier1 GEMM quick tuning space were exposed, the worst of them by a factor of 75.
Technical Details
makeGroupedGridLayoutderivesm_blockfrom the grid-wide block id asbid % thisMBlocksPerGroup. That divisor is a runtime value, soAMDGPUCodeGenPreparereplaced the remainder with its float-based algorithm once it proved both operands fit in 24 bits.That algorithm computes
trunc(fa * rcp(fb))and corrects with|fma(-fq, fb, fa)| >= fb, which only compensates for an underestimated quotient. An overestimate slips through unchanged. Past roughly 2^23 the 1 ulp error ofv_rcp_f32is enough to push the product over an integer boundary, so the quotient comes out one too large, the remainder goes negative, is masked back to 24 bits, andm_blocklands far outside the tensor. Nothing clamps the access, because buffer descriptors setnum_recordsto0x7ffffffeso that the0x80000000disabled-lane sentinel still falls out of range.Whether a given group size trips this depends on its reciprocal. The heuristic picks 11 on gfx90a, and
v_rcp_f32(11)is correctly rounded, so 237056 of the 14155776 workgroups addressed memory about 24 MB past the end of A. gfx942 escaped only by luck: its group size of 7 has a reciprocal one ulp low, which cannot overestimate.This is a backend bug rather than a grid-layout one, and it is already fixed upstream, so this cherry-picks both parts instead of working around it. Only llvm/llvm-project#201186 is strictly needed for the site above, since its dividend needs 24 bits and that patch rejects anything above 23. It leaves a hole though: the earliest dividend where the old expansion goes wrong is 8290592, a 23-bit value that would still take the float path. llvm/llvm-project#202753 closes it by dropping the unsigned bound to 22 bits, and it does not apply on its own since it builds on the first. Both are recorded under
llvm-patches/following the convention for downstream cherry-picks.Taking the fix in the backend also covers every other div/rem site with a non-constant divisor, not just this one.
makeGxNGridLayouthasbid / (splitKV * mBlocks)in the attention path, which was never audited.Test Plan
[0, 14155776), cross-checked against a CPU model, plus an exhaustive search for the smallest dividend at which it first becomes inexact over all divisors below 2^24.mlir/test/fusion/pr-e2e/rock-gemm-large-grid-float-divrem.mlir, the failing GEMM withgridGroupSize=11pinned so it does not depend on the CU count.Test Result
GridLayoutEmitter.cppuntouched, the grid-widebiddivision now goes through the exact 32-bit expansion, a reciprocal estimate followed by Newton refinement and an exact correction, while the group-local division stays on the cheap float path because its operand is masked to 16 bits. The backend now separates the safe and unsafe cases on its own.GridLayoutEmitter.cppuntouched.git clang-formatreports no changes.Submission Checklist