[AIROCMLIR-899] Enable weekly CI stages (parameter sweeps + tuning) - #361
bogdan-petkovic wants to merge 40 commits into
Conversation
Signed-off-by: bogdan-petkovic <bpetkovi@amd.com>
Signed-off-by: bogdan-petkovic <bpetkovi@amd.com>
…OCMLIR-899-enable-weekly-ci # Conflicts: # mlir/utils/jenkins/Jenkinsfile
8d87eeb to
ec1568d
Compare
…OCMLIR-899-enable-weekly-ci # Conflicts: # mlir/utils/jenkins/Jenkinsfile
There was a problem hiding this comment.
Pull request overview
This PR enables the previously commented-out weekly Jenkins CI stages in rocmlirTriton (parameter sweeps, kernel tuning, and perfDB archival) so weekly jobs actually execute and publish tuning artifacts.
Changes:
- Enables weekly-stage helper functions and stage blocks (while keeping the nightly benchmark/code-coverage blocks commented out).
- Adds/reenables Jenkins parameters that gate weekly execution (
weeklyTasks,sharedLib,staticLib) and skips the normal Build-and-Test matrix on weekly runs. - Adjusts weekly-stage execution to use
bash cmake.shas the rocmlirTriton build entrypoint and publishes per-arch tuning DB artifacts earlier.
Suppressed comments (1)
mlir/utils/jenkins/Jenkinsfile:1587
- The later tuning steps also pipe
tuningRunner.pyoutput intotee, so the Python process can fail without failing theshstep (pipeline exit code comes fromteeunlesspipefailis set). This can make the subsequent log-scan the only failure signal and miss cases where the process exits non-zero without emitting an "error" line. Prefer running these piped commands withbash -o pipefail(orset -o pipefail) so the stage fails on the real exit code.
// Find errors that are not part of a warning line
def errors = tuneLog.findAll { it =~ /(?i)error/ && !(it =~ /(?i)\bWARNING\b.*error/) }
if (errors) {
currentBuild.result = 'FAILURE'
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Verdict: REQUEST_CHANGES -- submitted as COMMENT (automated reviews are advisory) · Findings: 1 (0 Critical, 1 Major, 0 Minor)
Scope
Single-file change to mlir/utils/jenkins/Jenkinsfile that turns on the previously block-commented weekly CI stages (Parameter sweeps, Tune MLIR kernels, Archive weekly tuning perfDB), uncomments the params and helpers they depend on (sharedLib, staticLib, weeklyTasks, getLabelFromChip, shouldRunFromChip, archivePerfDB, runSweep), enables the Build and Test weekly guard and the gfx950 8-GPU weekly node label, swaps buildProject(...) for bash cmake.sh, and adds a per-branch archiveArtifacts of build/*.tsv.
Findings
One Major finding at mlir/utils/jenkins/Jenkinsfile:1584 — the tuning-log grep is the only failure signal for most tuningRunner.py invocations, because the | tee pipelines discard the exit code and no pipefail is set. See the inline comment.
Notes
Spot-checks that came back clean, recorded so they don't get re-litigated:
- The
buildProject('check-rocmlir-build-only ci-performance-scripts', '')->bash cmake.shswap does not drop the perf-script bundle:cmake.sh:25runsninja check-rocmlir-build-only, andci-performance-scriptsis already listed inROCMLIR_TEST_DEPENDS(mlir/test/CMakeLists.txt:102), which that target depends on. Sobuild/bin/{parameterSweeps,attentionSweeps,tuningRunner}.pyandrocm-runare still staged. - The new
--test-timeout-sec 600inrunSweepis valid for both sweep scripts: it is registered inparameterSweeps.py:1108(default 120, matching the comment) andattentionSweeps.pyimports the sameadd_common_args. params.weeklyis defined (Jenkinsfile:1116), so the newly enabledBuild and Testguard at line 1185 is safe; the Jenkins stage view on this PR confirms the file parses andBuild and Teststill ran.- Per-branch
archiveArtifactsin thefinallyblock is a genuine fix:archivePerfDB()usesonlyIfSuccessful: trueandparallelsAlwaysFailFast()is on (line 1109), so a single failing chip would otherwise lose every branch's tuning DB.
Two things worth confirming before merge, neither anchorable to a diff hunk:
- The PR description says
Tune Fusionnow uses "the canonical--test-dirflag", but lines 1601 and 1604 still pass--test_dir. The repo does register underscore aliases elsewhere (--rocmlir-gen-flags/--rocmlir_gen_flagsintuningArgumentUtils.py:73), so this may be fine — please confirmtuningRunner.pyaccepts--test_dir, or update the description. - The description also still claims the tuning commands run through
shStrict; commit911237b3creplaced that with plainsh. Worth refreshing the description so it matches the head commit.
No rocMLIR back-port note is required: mlir/utils/jenkins/Jenkinsfile is not in the shared-with-rocMLIR path list (only mlir/utils/jenkins/static-checks/ is), and this file has already diverged by having these stages commented out downstream.
CI status
No failing checks. Jenkins, review, and copilot-pull-request-reviewer are in progress; the weekly stages report PENDING because this is not a weekly run.
… use canonical --test-dir
…OCMLIR-899-enable-weekly-ci Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # mlir/utils/performance/perfRunner.py # mlir/utils/performance/tuningRunner.py
…, use it in weekly Co-authored-by: Cursor <cursoragent@cursor.com>
…ning OOM Co-authored-by: Cursor <cursoragent@cursor.com>
…tead of aborting Co-authored-by: Cursor <cursoragent@cursor.com>
The stage inherited tuningRunner's `full` default. A full run measured 39.6h on gfx90a for its five commands, with conv alone taking 25.7h, and did not finish at all on the other five chips. Expose the space as a `weeklyTuningSpace` choice defaulting to `quick`, which the same run measured 7 to 10 times faster per command, and keep `full` and `exhaustive` for a deliberate re-tune. The separate quick-DB commands now only run when the stage itself tuned a wider space, since otherwise they recompute what the main commands just did. Also drop the measurement-only knobs. `--gpu-run-timeout 300` treated a legitimately slow config as a hang and lost the whole problem, and `--num-cpus 8` bounded attention compiles only to survive that run, so `--abort-on-error` is restored in place of collecting failures. The verification bound stays disabled because verifying the largest tier1 gemm was measured at 104 minutes against the 600s default.
…OCMLIR-899-enable-weekly-ci
…n path The six rocmlir-driver invocations in perfRunner.py's fusion path still pass -targets, which came in with the initial import from rocMLIR and was orphaned when #80 restructured the driver's pipelines and dropped that option. rocmlir-driver rejects the flag, the stage produces no output, and because the call sends stderr to DEVNULL the failure surfaces much later as an IndexError on result[3] in get_fusion_test_info. Nothing exercised this until now. The fusion benchmarking calls sit in the commented-out nightly stage, and the fusion tuning calls only run from the weekly. In rocMLIR -targets takes a comma-separated list tied to the MHAL multi-target path while -arch takes a single target, and this code always passes one chip, so -arch is the faithful replacement. Verified by extracting the tuning key for every file under mlir/test/fusion/resnet50-e2e and mlir/test/xmir/bert-torch-tosa-e2e on gfx90a and gfx1201, 50 files in total, all of which now yield a test vector.
format_error bounds reported output at ten lines, keeping the first and last five, which is right for a failed verification but wrong for a tuning driver that dies on a signal. The weekly hit an assertion in Triton's PaddedSharedEncodingAttr on gfx950 three runs in a row and the report elided the middle 39 frames, which are the ones naming the pass that constructs the attribute, leaving the failure undiagnosable. Raise the bound for this call site only. Verified that a 50 line stderr now survives whole instead of collapsing to 13 lines with an elision marker.
8d516df to
d15bd6e
Compare
| Value n_block = | ||
| DivUIOp::create(b, loc, RemUIOp::create(b, loc, bid, blocksPerGroup), | ||
| thisMBlocksPerGroup); | ||
| // Index within the group rather than the grid: the grid-wide bid can exceed |
There was a problem hiding this comment.
this is not related to enabling weekly CI. It needs to be an independent PR
d15bd6e to
42afd25
Compare
…OCMLIR-899-enable-weekly-ci
42afd25 to
06c4c3a
Compare
…OCMLIR-899-enable-weekly-ci
#559) * Skip gfx950 attention perf configs that LLVM miscompiles in the sweeps * Check SKIPPED_PERF_CONFIGS keys against PERF_CONFIG_FIELD_NAMES at import
…OCMLIR-899-enable-weekly-ci
…OCMLIR-899-enable-weekly-ci
Motivation
resolve: https://amd-hub.atlassian.net/browse/AIROCMLIR-899
Enable the weekly CI stages for rocmlirTriton. The weekly pipeline (parameter sweeps + kernel tuning + weekly tuning-DB archival) already existed in
mlir/utils/jenkins/Jenkinsfilebut was fully commented out, so weekly CI never actually ran. This PR turns those stages on and adapts them to the rocmlirTriton build.Technical Details
All changes are in
mlir/utils/jenkins/Jenkinsfile:Parameter sweeps,Tune MLIR kernels, andArchive weekly tuning perfDBare live, whileBenchmark and Report PerformanceandCode coverageremain commented out.when {}guards depend on:sharedLib,staticLib, andweeklyTasks(default/parameterSweeps/Tuning).Build and Testwhen { equals expected: false, actual: params.weekly }guard so the PR/nightly build-and-test stage is skipped on weekly runs.getLabelFromChip,shouldRunFromChip,archivePerfDB, andrunSweep(the still-unusedbuild_fixedE2ETests,check_randomE2ETests,isNotNavi3x,collectCoverageData, andpreMergeCheckPackagehelpers stay commented).gfx950weekly node label so weekly runs usemlir && linux-mi350-8(8-GPU) instead of the single-GPUmlir && linux-mi350-1.buildProject('check-rocmlir-build-only ci-performance-scripts', '')with the rocmlirTriton build entrypoint (bash cmake.shunder a 120-minute timeout) in the weeklyPrepare Performance ScriptsandTune rocMLIRstages —buildProjectis commented out ondevelop, so the stages would not otherwise compile.Tune rocMLIRstage: run the tuning commands throughshStrictso a failingtuningRunner.pyin a| teepipeline is no longer masked by tee's exit code; filter outerrormatches that are part of aWARNINGline to avoid false failures; report the matched error lines and setcurrentBuild.result = 'FAILURE'before callingerror(...); and use the canonical--test-dirflag inTune Fusion.build/*.tsvin each tuning branch'sfinallyblock (allowEmptyArchive: true,onlyIfSuccessful: false) so per-arch tuning databases are published as soon as a branch finishes, without waiting for the other parallel matrix branches.Test Plan
weeklyTasksselector (parameterSweepsandTuning) and verify only the corresponding stages execute.Build and Testruns as before).Test Result
Submission Checklist