Skip to content

[AIROCMLIR-1318] Fix perf benchmark tooling: hipBLASLt/CK baselines, fusion benchmarks, fusion report - #545

Merged
Muhamed-Husic merged 8 commits into
developfrom
mhusic/AIROCMLIR-1318-tooling-fixes
Oct 2, 2026
Merged

Muhamed-Husic merged 8 commits into
developfrom
mhusic/AIROCMLIR-1318-tooling-fixes

Conversation

@Muhamed-Husic

Copy link
Copy Markdown
Collaborator

Motivation

While enabling the nightly "Benchmark and Report Performance" stage in rocmlirTriton (AIROCMLIR-901), I ran the stage's perfRunner.py and report commands locally on gfx942 and hit three issues in the perf tooling. The stage hasn't been running in rocmlirTriton so far, so nothing exercised these paths. All three are present in current develop:

  1. hipBLASLt and CK baselines are always empty. Add -transO to GEMM tuning problem config #284 added -transO to the GEMM problem config, and perfRunner passes it to hipblaslt-benchmark-driver and ck-gemm-benchmark-driver as well. Their shared argument parser doesn't recognize it yet and exits with Invalid argument!. perfRunner records that as NaN and carries on, so the "MLIR vs hipBLASLt" and "MLIR vs CK" reports would have no baseline.
  2. Every fusion benchmark fails. perfRunner compiles fused kernels with -fut=<name>_wrapper, but since [AIROCMLIR-548] Fix clone-harness behavior in rocmlir-gen #76 rocmlir-gen --clone-harness no longer creates a _wrapper function. With [AIROCMLIR-899] Enable weekly CI stages (parameter sweeps + tuning) #361's -arch fix, all 34 files in resnet50-e2e and bert-torch-tosa-e2e hit the "does -fut point to the wrong function?" assertion, and the Fusion TFlops column stays empty.
  3. The fusion report crashes. A few test files run the same conv or GEMM with different fused operations (mixr-resnet-fusion-case-1-quantization / -case-1-int8, bert_part_0 / bert_part_5). Each gets its own row (our tuning keys include the fused operations), but the report tells rows apart only by the conv/GEMM parameters, so it sees duplicates and fails ("non-unique index"). This happens whether or not the fusion runs succeed, so with the nightly stage on, "Create performance reports" would fail every night.

Technical Details

One commit per fix.

  1. Accept -transO in the benchmark drivers: the shared parser (common/benchmarkUtils.cpp) now accepts -transO= the same way as -transA=/-transB=. The hipBLASLt and CK drivers fail with a clear error if it's true, since neither driver handles transposed output (no tier1 gemm config uses -transO true).

  2. Use -fut=<name> for fused kernels: one-line change in benchmark_fusion_kernels(), matching the RUN lines of the fusion tests. It's also covered by a new test in perfRunner-test.py.

  3. Add FileName to the fusion report index: clean_data_for_humans() adds FileName to the row index when the column exists, and only fusion results have it. The colliding rows stay separate, since they're different kernels. The CSVs and all other reports are unchanged; in the fusion HTML reports, FileName just moves into the row labels. This is covered by a new test perf-scripts/fusion-performance-report.py

Dependencies

Test Plan

On gfx942 (MI300); rocmlirTriton built with -DROCMLIR_ENABLE_BENCHMARKS=hipblaslt:

  • perfRunner.py --op=gemm --batch-all on a sample of tier1 gemm configs, to check the hipBLASLt baseline.
  • perfRunner.py --op=fusion on resnet50-e2e and bert-torch-tosa-e2e, then createFusionPerformanceReports.py (with PR#361's -arch change applied locally, see Dependencies).
  • The new tests against the old and the fixed scripts, and LIT_FILTER=perf-scripts ninja check-rocmlir.

Test Result

  • The hipBLASLt baseline is filled in; before the fix the column was empty.
  • All 34 fusion files produce fusion TFlops, and both fusion reports are generated.
  • The new tests fail before the fixes and pass after; perf-scripts lit passes (31 tests).
  • The -transO check in ck-gemm-benchmark-driver.cpp isn't compiled yet, since Composable Kernel isn't installed locally and PR CI doesn't build the benchmark drivers It will first be built by the CK step of the nightly Benchmark stage added in AIROCMLIR-901.

Submission Checklist

rocmlir_gen_args = [
'-ph', '-fut=' + fut_name + '_wrapper', '--perf_config=' + best_perf, '-'
]
rocmlir_gen_args = ['-ph', '-fut=' + fut_name, '--perf_config=' + best_perf, '-']

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rocMLIR back-port check: every file this PR touches is under mlir/utils/performance/, which docs/PR_REVIEW_CHECKLIST.md ("rocMLIR back-port check -- Likely needs a back-port") lists as shared with ROCm/rocMLIR. The shared paths here are perfRunner.py, reportUtils.py, common/benchmarkUtils.{h,cpp}, hipblaslt-benchmark-driver/hipblaslt-benchmark-driver.cpp and ck-benchmark-driver/ck-gemm-benchmark-driver.cpp. The description's Dependencies section only refers to rocmlirTriton PR #361, so none of options (a)/(b)/(c) is satisfied. Please add either a link to a parallel rocMLIR PR, or a one-line note per fix explaining why it does not apply upstream -- fix 2 plausibly does not (the _wrapper removal came from this tree's #76), but the -transO parser fix and the fusion-report index fix look like they would apply to rocMLIR unchanged.


CONV_CSV = """\
Direction,DataType,Chip,numCU,numChiplets,FilterLayout,InputLayout,OutputLayout,N,C,H,W,K,Y,X,DilationH,DilationW,StrideH,StrideW,PaddingH,PaddingW,PerfConfig,LDSBankConflict,Fusion TFlops,MLIR TFlops,Fusion/MLIR,FileName
fwd,i8,gfx942,304,8,gkc01,ngc01,ngk01,1,128,56,56,128,3,3,1,1,2,2,1,1,,NaN,0.712234,1.522267,0.467877,mixr-resnet-fusion-case-1-quantization.mlir

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fixture's FileName column holds bare basenames, but benchmark_fusion_kernels() populates it from glob.glob(test_dir + '/*.mlir') (perfRunner.py:3048, assigned at :3101 and :3119), so the real CSV always carries a test_dir-prefixed path. The docstring claims these rows "are taken from a gfx942 run over resnet50-e2e and bert-torch-tosa-e2e", which is then inaccurate. The test still exercises the duplicate-index failure either way, but please either use the path form perfRunner actually writes so the fixture matches production data, or drop the provenance claim from the docstring.

import unittest
from pathlib import Path

PERF_DIR = Path(__file__).resolve().parents[2] / "utils" / "performance"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This imports createFusionPerformanceReports and reportUtils from the source tree, while the sibling perfRunner-test.py:24-32 in the same directory deliberately resolves the script via shutil.which() and documents why (the deployed copies under ROCMLIR_BIN_DIR are what ci-performance-scripts ships and what the Jenkins perf stage runs). Neither module needs the compiled amd_arch_db binding, so the source-tree import works today, but it leaves two opposing conventions side by side and means this test would still pass if a script were dropped from PERFORMANCE_SCRIPTS in mlir/utils/performance/CMakeLists.txt. Prefer resolving the path the same way the sibling test does, or add a short comment stating why this file intentionally differs.

@rocmlir-pr-reviewer rocmlir-pr-reviewer Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: COMMENT  ·  Findings: 3 (0 Critical, 1 Major, 2 Minor)


Scope

Three independent fixes to the performance tooling, one commit each: (1) -transO= is now accepted by the shared benchmark-driver argument parser (common/benchmarkUtils.{h,cpp}), with the hipBLASLt and CK drivers rejecting true explicitly; (2) benchmark_fusion_kernels() times the fused kernel under -fut=<name> instead of the no-longer-generated <name>_wrapper; (3) clean_data_for_humans() adds FileName to the row index so two files fusing different ops around the same problem no longer collide into a non-unique index. Two new pure-Python lit tests cover (2) and (3).

Findings

  • mlir/utils/performance/perfRunner.py:3106 — every changed source file lives under mlir/utils/performance/, which the checklist lists as shared with ROCm/rocMLIR; the PR description has no back-port note (Major).
  • mlir/test/perf-scripts/fusion-performance-report.py:34 — the fixture's FileName values are bare basenames, but benchmark_fusion_kernels() writes the glob path (Minor).
  • mlir/test/perf-scripts/fusion-performance-report.py:23 — imports the scripts from the source tree while the sibling test in the same directory deliberately imports from PATH (Minor).

Notes

Spot-checks that came out clean:

  • The -transO= branch in parseCommandLine() matches the existing -transA=/-transB= idiom exactly, and generate_problem_commandline() emits -transO=True/-transO=False, which atob() handles. printUsage() and printProblem() were both updated.
  • run_fusion_kernel() already invokes rocmlir-gen -fut <fut_name> --clone-harness, so the new -fut=<name> in the second rocmlir-gen stage is consistent with the harness it consumes.
  • FileName is only ever written by benchmark_fusion_kernels(), so the reportUtils index change cannot reach the MLIR-vs-hipBLASLt/CK/MIOpen or regression reports. It is appended after PerfConfig, but nothing consumes the returned index_cols positionally (perfRegressionReport.py slices the raw *_TEST_PARAMETERS lists, not index_cols).
  • Both new tests genuinely fail without their fix: assertIn over the argv list is an exact-element match, so _wrapper would be caught, and the duplicate-index rows make pandas' Styler raise.

Out of scope: the -transO handling in the CK/hipBLASLt drivers has no automated coverage because those drivers are not built in PR CI, as the PR description acknowledges.

CI status

No failing or cancelled checks. Jenkins, Build and Test, MIGraphX and Code coverage are still pending; py-checks passed.

@bogdan-petkovic bogdan-petkovic left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@Muhamed-Husic
Muhamed-Husic enabled auto-merge (squash) October 2, 2026 10:17
@Muhamed-Husic
Muhamed-Husic merged commit 9f1ee60 into develop Oct 2, 2026
6 of 9 checks passed
@Muhamed-Husic
Muhamed-Husic deleted the mhusic/AIROCMLIR-1318-tooling-fixes branch October 2, 2026 12:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants