Skip to content

[AIROCMLIR-899] Enable weekly CI stages (parameter sweeps + tuning) - #361

Open
bogdan-petkovic wants to merge 40 commits into
developfrom
users/bpetkovi/AIROCMLIR-899-enable-weekly-ci
Open

bogdan-petkovic wants to merge 40 commits into
developfrom
users/bpetkovi/AIROCMLIR-899-enable-weekly-ci

Conversation

@bogdan-petkovic

Copy link
Copy Markdown
Collaborator

Motivation

resolve: https://amd-hub.atlassian.net/browse/AIROCMLIR-899

Enable the weekly CI stages for rocmlirTriton. The weekly pipeline (parameter sweeps + kernel tuning + weekly tuning-DB archival) already existed in mlir/utils/jenkins/Jenkinsfile but was fully commented out, so weekly CI never actually ran. This PR turns those stages on and adapts them to the rocmlirTriton build.

Technical Details

All changes are in mlir/utils/jenkins/Jenkinsfile:

  • Enable the three weekly stages by moving the block-comment boundary so Parameter sweeps, Tune MLIR kernels, and Archive weekly tuning perfDB are live, while Benchmark and Report Performance and Code coverage remain commented out.
  • Uncomment the pipeline parameters the weekly when {} guards depend on: sharedLib, staticLib, and weeklyTasks (default / parameterSweeps / Tuning).
  • Enable the Build and Test when { equals expected: false, actual: params.weekly } guard so the PR/nightly build-and-test stage is skipped on weekly runs.
  • Uncomment the helper functions the weekly stages call: getLabelFromChip, shouldRunFromChip, archivePerfDB, and runSweep (the still-unused build_fixedE2ETests, check_randomE2ETests, isNotNavi3x, collectCoverageData, and preMergeCheckPackage helpers stay commented).
  • Enable the gfx950 weekly node label so weekly runs use mlir && linux-mi350-8 (8-GPU) instead of the single-GPU mlir && linux-mi350-1.
  • Replace buildProject('check-rocmlir-build-only ci-performance-scripts', '') with the rocmlirTriton build entrypoint (bash cmake.sh under a 120-minute timeout) in the weekly Prepare Performance Scripts and Tune rocMLIR stages — buildProject is commented out on develop, so the stages would not otherwise compile.
  • Harden the Tune rocMLIR stage: run the tuning commands through shStrict so a failing tuningRunner.py in a | tee pipeline is no longer masked by tee's exit code; filter out error matches that are part of a WARNING line to avoid false failures; report the matched error lines and set currentBuild.result = 'FAILURE' before calling error(...); and use the canonical --test-dir flag in Tune Fusion.
  • Archive build/*.tsv in each tuning branch's finally block (allowEmptyArchive: true, onlyIfSuccessful: false) so per-arch tuning databases are published as soon as a branch finishes, without waiting for the other parallel matrix branches.

Test Plan

  • Weekly CI
  • Exercise the weeklyTasks selector (parameterSweeps and Tuning) and verify only the corresponding stages execute.
  • Confirm a regular PR CI run is unaffected (weekly stages skipped, Build and Test runs as before).

Test Result

  • Weekly CI

Submission Checklist

Signed-off-by: bogdan-petkovic <bpetkovi@amd.com>
Signed-off-by: bogdan-petkovic <bpetkovi@amd.com>
…OCMLIR-899-enable-weekly-ci

# Conflicts:
#	mlir/utils/jenkins/Jenkinsfile
@bogdan-petkovic
bogdan-petkovic force-pushed the users/bpetkovi/AIROCMLIR-899-enable-weekly-ci branch from 8d87eeb to ec1568d Compare August 5, 2026 13:15
@bogdan-petkovic
bogdan-petkovic marked this pull request as ready for review August 7, 2026 10:16
@bogdan-petkovic bogdan-petkovic added the claude-review Request a Claude PR review label Aug 7, 2026
@bogdan-petkovic
bogdan-petkovic requested review from mirza-halilcevic and a lite review from Copilot August 7, 2026 10:42

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR enables the previously commented-out weekly Jenkins CI stages in rocmlirTriton (parameter sweeps, kernel tuning, and perfDB archival) so weekly jobs actually execute and publish tuning artifacts.

Changes:

  • Enables weekly-stage helper functions and stage blocks (while keeping the nightly benchmark/code-coverage blocks commented out).
  • Adds/reenables Jenkins parameters that gate weekly execution (weeklyTasks, sharedLib, staticLib) and skips the normal Build-and-Test matrix on weekly runs.
  • Adjusts weekly-stage execution to use bash cmake.sh as the rocmlirTriton build entrypoint and publishes per-arch tuning DB artifacts earlier.
Suppressed comments (1)

mlir/utils/jenkins/Jenkinsfile:1587

  • The later tuning steps also pipe tuningRunner.py output into tee, so the Python process can fail without failing the sh step (pipeline exit code comes from tee unless pipefail is set). This can make the subsequent log-scan the only failure signal and miss cases where the process exits non-zero without emitting an "error" line. Prefer running these piped commands with bash -o pipefail (or set -o pipefail) so the stage fails on the real exit code.
                                                                // Find errors that are not part of a warning line
                                                                def errors = tuneLog.findAll { it =~ /(?i)error/ && !(it =~ /(?i)\bWARNING\b.*error/) }

                                                                if (errors) {
                                                                    currentBuild.result = 'FAILURE'

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread mlir/utils/jenkins/Jenkinsfile Outdated
Comment thread mlir/utils/jenkins/Jenkinsfile
Comment thread mlir/utils/jenkins/Jenkinsfile

@rocmlir-pr-reviewer rocmlir-pr-reviewer Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: REQUEST_CHANGES -- submitted as COMMENT (automated reviews are advisory)  ·  Findings: 1 (0 Critical, 1 Major, 0 Minor)


Scope

Single-file change to mlir/utils/jenkins/Jenkinsfile that turns on the previously block-commented weekly CI stages (Parameter sweeps, Tune MLIR kernels, Archive weekly tuning perfDB), uncomments the params and helpers they depend on (sharedLib, staticLib, weeklyTasks, getLabelFromChip, shouldRunFromChip, archivePerfDB, runSweep), enables the Build and Test weekly guard and the gfx950 8-GPU weekly node label, swaps buildProject(...) for bash cmake.sh, and adds a per-branch archiveArtifacts of build/*.tsv.

Findings

One Major finding at mlir/utils/jenkins/Jenkinsfile:1584 — the tuning-log grep is the only failure signal for most tuningRunner.py invocations, because the | tee pipelines discard the exit code and no pipefail is set. See the inline comment.

Notes

Spot-checks that came back clean, recorded so they don't get re-litigated:

  • The buildProject('check-rocmlir-build-only ci-performance-scripts', '') -> bash cmake.sh swap does not drop the perf-script bundle: cmake.sh:25 runs ninja check-rocmlir-build-only, and ci-performance-scripts is already listed in ROCMLIR_TEST_DEPENDS (mlir/test/CMakeLists.txt:102), which that target depends on. So build/bin/{parameterSweeps,attentionSweeps,tuningRunner}.py and rocm-run are still staged.
  • The new --test-timeout-sec 600 in runSweep is valid for both sweep scripts: it is registered in parameterSweeps.py:1108 (default 120, matching the comment) and attentionSweeps.py imports the same add_common_args.
  • params.weekly is defined (Jenkinsfile:1116), so the newly enabled Build and Test guard at line 1185 is safe; the Jenkins stage view on this PR confirms the file parses and Build and Test still ran.
  • Per-branch archiveArtifacts in the finally block is a genuine fix: archivePerfDB() uses onlyIfSuccessful: true and parallelsAlwaysFailFast() is on (line 1109), so a single failing chip would otherwise lose every branch's tuning DB.

Two things worth confirming before merge, neither anchorable to a diff hunk:

  • The PR description says Tune Fusion now uses "the canonical --test-dir flag", but lines 1601 and 1604 still pass --test_dir. The repo does register underscore aliases elsewhere (--rocmlir-gen-flags / --rocmlir_gen_flags in tuningArgumentUtils.py:73), so this may be fine — please confirm tuningRunner.py accepts --test_dir, or update the description.
  • The description also still claims the tuning commands run through shStrict; commit 911237b3c replaced that with plain sh. Worth refreshing the description so it matches the head commit.

No rocMLIR back-port note is required: mlir/utils/jenkins/Jenkinsfile is not in the shared-with-rocMLIR path list (only mlir/utils/jenkins/static-checks/ is), and this file has already diverged by having these stages commented out downstream.

CI status

No failing checks. Jenkins, review, and copilot-pull-request-reviewer are in progress; the weekly stages report PENDING because this is not a weekly run.

@rocmlir-pr-reviewer rocmlir-pr-reviewer Bot removed the claude-review Request a Claude PR review label Aug 7, 2026
@bogdan-petkovic
bogdan-petkovic requested a balanced review from Copilot August 7, 2026 11:51
@bogdan-petkovic bogdan-petkovic added the claude-review Request a Claude PR review label Aug 7, 2026
bogdan-petkovic and others added 10 commits September 8, 2026 06:34
…OCMLIR-899-enable-weekly-ci

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	mlir/utils/performance/perfRunner.py
#	mlir/utils/performance/tuningRunner.py
…, use it in weekly

Co-authored-by: Cursor <cursoragent@cursor.com>
…ning OOM

Co-authored-by: Cursor <cursoragent@cursor.com>
…tead of aborting

Co-authored-by: Cursor <cursoragent@cursor.com>
The stage inherited tuningRunner's `full` default. A full run measured
39.6h on gfx90a for its five commands, with conv alone taking 25.7h,
and did not finish at all on the other five chips.

Expose the space as a `weeklyTuningSpace` choice defaulting to `quick`,
which the same run measured 7 to 10 times faster per command, and keep
`full` and `exhaustive` for a deliberate re-tune. The separate quick-DB
commands now only run when the stage itself tuned a wider space, since
otherwise they recompute what the main commands just did.

Also drop the measurement-only knobs. `--gpu-run-timeout 300` treated a
legitimately slow config as a hang and lost the whole problem, and
`--num-cpus 8` bounded attention compiles only to survive that run, so
`--abort-on-error` is restored in place of collecting failures. The
verification bound stays disabled because verifying the largest tier1
gemm was measured at 104 minutes against the 600s default.
…n path

The six rocmlir-driver invocations in perfRunner.py's fusion path still
pass -targets, which came in with the initial import from rocMLIR and
was orphaned when #80 restructured the driver's pipelines and dropped
that option. rocmlir-driver rejects the flag, the stage produces no
output, and because the call sends stderr to DEVNULL the failure
surfaces much later as an IndexError on result[3] in
get_fusion_test_info.

Nothing exercised this until now. The fusion benchmarking calls sit in
the commented-out nightly stage, and the fusion tuning calls only run
from the weekly. In rocMLIR -targets takes a comma-separated list tied
to the MHAL multi-target path while -arch takes a single target, and
this code always passes one chip, so -arch is the faithful replacement.

Verified by extracting the tuning key for every file under
mlir/test/fusion/resnet50-e2e and mlir/test/xmir/bert-torch-tosa-e2e on
gfx90a and gfx1201, 50 files in total, all of which now yield a test
vector.
format_error bounds reported output at ten lines, keeping the first and
last five, which is right for a failed verification but wrong for a
tuning driver that dies on a signal. The weekly hit an assertion in
Triton's PaddedSharedEncodingAttr on gfx950 three runs in a row and the
report elided the middle 39 frames, which are the ones naming the pass
that constructs the attribute, leaving the failure undiagnosable.

Raise the bound for this call site only. Verified that a 50 line stderr
now survives whole instead of collapsing to 13 lines with an elision
marker.
Value n_block =
DivUIOp::create(b, loc, RemUIOp::create(b, loc, bid, blocksPerGroup),
thisMBlocksPerGroup);
// Index within the group rather than the grid: the grid-wide bid can exceed

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is not related to enabling weekly CI. It needs to be an independent PR

@bogdan-petkovic
bogdan-petkovic force-pushed the users/bpetkovi/AIROCMLIR-899-enable-weekly-ci branch from d15bd6e to 42afd25 Compare September 28, 2026 18:31
@bogdan-petkovic
bogdan-petkovic force-pushed the users/bpetkovi/AIROCMLIR-899-enable-weekly-ci branch from 42afd25 to 06c4c3a Compare September 29, 2026 18:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants