Repository navigation
ci: open an issue when a run nobody is watching fails #1146
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,84 @@ | ||
| name: "Report an unattended failure (reusable)" | ||
|
|
||
| # Records a failed run that nobody is watching on a GitHub issue. Called as the | ||
| # last job of fuzz.yml, bench-eql.yml, macro-expand-eql.yml and test-eql.yml. | ||
| # | ||
| # WHY. A scheduled run, or a push to main, has no pull request to turn red. | ||
| # GitHub mails a scheduled run's failure to one person — whoever last edited | ||
| # the cron line — and a push run's to whoever pushed. Everyone else learns of | ||
| # it only by opening the Actions tab. An issue notifies the repository's | ||
| # watchers and stays open until someone deals with it. | ||
| # | ||
| # ONE ISSUE PER WORKFLOW AND BRANCH, NOT PER RUN. The issue is found by its | ||
| # exact title, which names both. If one is open, the failure is added to it as | ||
| # a comment, so a workflow that keeps failing nightly builds one thread rather | ||
| # than a pile of duplicates. Once it is closed, the next failure opens a new | ||
| # one. | ||
| # | ||
| # The caller decides when this runs and grants the permission, because a called | ||
| # workflow cannot raise its caller's token: | ||
| # | ||
| # report: | ||
| # needs: [<every other job>] | ||
| # if: >- | ||
| # failure() && github.event_name != 'pull_request' && | ||
| # github.event_name != 'workflow_dispatch' | ||
| # permissions: | ||
| # issues: write | ||
| # uses: ./.github/workflows/_report-unattended-failure.yml | ||
| # with: | ||
| # guidance: <what the person who picks this up should do> | ||
| # | ||
| # A pull request shows its own failure, and whoever dispatches a run by hand is | ||
| # watching it (and may have pointed it at any branch). Every other event is | ||
| # reported, so a trigger added later is covered without editing the condition. | ||
| # | ||
| # `needs` must list every other job in the workflow: `failure()` only sees the | ||
| # jobs a job needs, so a failure in one left out is not reported. | ||
| # scripts/__tests__/unattended-failure-report.test.mjs checks this, and that | ||
| # every scheduled workflow either calls this file or is listed there with the | ||
| # reason it does not. | ||
| on: | ||
| workflow_call: | ||
| inputs: | ||
| guidance: | ||
| description: "What to do about the failure. Goes in the issue body and in every comment." | ||
| required: true | ||
| type: string | ||
|
|
||
| permissions: | ||
| contents: read | ||
|
|
||
| jobs: | ||
| report: | ||
| name: Open or update the failure issue | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 5 | ||
| permissions: | ||
| issues: write | ||
| steps: | ||
| - name: Open or update the failure issue | ||
| env: | ||
| GH_TOKEN: ${{ github.token }} | ||
| GH_REPO: ${{ github.repository }} | ||
| # `github.workflow` in a called workflow is the CALLER's name. | ||
| TITLE: "${{ github.workflow }} failed on ${{ github.ref_name }}" | ||
| EVENT: ${{ github.event_name }} | ||
| SHA: ${{ github.sha }} | ||
| RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} | ||
| GUIDANCE: ${{ inputs.guidance }} | ||
| run: | | ||
| set -euo pipefail | ||
| body="The \`$EVENT\` run at $SHA failed: $RUN_URL | ||
|
|
||
| $GUIDANCE" | ||
| # `--search` matches words, not the whole title; the jq filter | ||
| # demands an exact match so a similarly named workflow's issue is | ||
| # never commented on. | ||
| issue=$(gh issue list --state open --search "in:title \"$TITLE\"" \ | ||
| --json number,title --jq "map(select(.title == env.TITLE)) | .[0].number // empty") | ||
| if [ -n "$issue" ]; then | ||
| gh issue comment "$issue" --body "$body" | ||
| else | ||
| gh issue create --title "$TITLE" --label needs-triage --body "$body" | ||
|
Comment on lines
+80
to
+83
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
When two push runs of the same workflow/ref fail concurrently (particularly the hour-long Useful? React with 👍 / 👎. |
||
| fi | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,125 @@ | ||
| import { describe, expect, it } from 'vitest' | ||
| import { expr } from './lib/expressions.mjs' | ||
| import { readWorkflow, workflowFiles } from './lib/workflows.mjs' | ||
|
|
||
| /** | ||
| * A scheduled workflow that fails must tell someone. | ||
| * | ||
| * WHAT HAPPENED. The nightly fuzz campaign uploaded each crash reproducer as an | ||
| * artifact and stopped there. A scheduled run has no pull request to turn red, | ||
| * and GitHub mails its failure to one person — whoever last edited the cron | ||
| * line — so a crash would have sat in the Actions tab until someone happened to | ||
| * look. test-eql.yml, macro-expand-eql.yml and bench-eql.yml had the same gap, | ||
| * as did the push-to-main runs of the first and last; only | ||
| * musl-build-image.yml opened an issue. | ||
| * | ||
| * THE FIX is `_report-unattended-failure.yml`, called as the last job of each | ||
| * of those workflows. This file pins the ways that call can quietly stop | ||
| * working: | ||
| * | ||
| * 1. A new scheduled workflow lands without it. So scheduled workflows are | ||
| * DISCOVERED by scanning the directory, and each must either call the | ||
| * reporter or be listed in `REPORTS_ELSEWHERE` with its reason. | ||
| * 2. A job is added to a workflow but not to the report job's `needs`. The | ||
| * report job's `failure()` only sees the jobs it needs, so a failure in the | ||
| * new job would go unreported, silently. So `needs` must be every other job. | ||
| * 3. The condition is narrowed back to `github.event_name == 'schedule'`, | ||
| * which drops push-to-main failures and fails shut on any trigger added | ||
| * later (eql-matrix-triggers.test.mjs forbids that shape in test-eql.yml | ||
| * for the same reason). So the condition is held to one spelling. | ||
| */ | ||
|
|
||
| const REPORTER = './.github/workflows/_report-unattended-failure.yml' | ||
|
|
||
| /** The report job's condition, as the parsed workflow holds it. */ | ||
| const CONDITION = | ||
| "failure() && github.event_name != 'pull_request' && github.event_name != 'workflow_dispatch'" | ||
|
|
||
| /** | ||
| * Scheduled workflows that do not call the reporter, with the reason. An | ||
| * equality, not a floor: an entry for a workflow that has since started | ||
| * calling the reporter, or stopped being scheduled, fails too. | ||
| */ | ||
| const REPORTS_ELSEWHERE = { | ||
| // CodeQL uploads its findings to code scanning, where they raise alerts on | ||
| // their own. | ||
| '.github/workflows/codeql.yml': 'findings go to code scanning', | ||
| // Same: `fail-on-vuln: false`, and the scan's SARIF goes to code scanning. | ||
| '.github/workflows/osv-scanner.yml': 'findings go to code scanning', | ||
| // Opens its own issue from a step inside its single job, pinned by | ||
| // musl-build-image.test.mjs. | ||
| '.github/workflows/musl-build-image.yml': 'opens its own issue', | ||
| } | ||
|
|
||
| /** | ||
| * The workflows known to call the reporter today. A minimum, so the discovery | ||
| * below cannot pass by finding nothing. | ||
| */ | ||
| const EXPECTED_REPORTERS = [ | ||
| '.github/workflows/bench-eql.yml', | ||
| '.github/workflows/fuzz.yml', | ||
| '.github/workflows/macro-expand-eql.yml', | ||
| '.github/workflows/test-eql.yml', | ||
| ] | ||
|
|
||
| function triggers(wf) { | ||
| return wf?.on ?? wf?.[true] ?? {} | ||
| } | ||
|
|
||
| const workflows = workflowFiles().map((path) => ({ | ||
| path, | ||
| wf: readWorkflow(path), | ||
| })) | ||
|
|
||
| function reportJobs(wf) { | ||
| return Object.entries(wf?.jobs ?? {}).filter( | ||
| ([, job]) => job?.uses === REPORTER, | ||
| ) | ||
| } | ||
|
|
||
| const reporters = workflows.filter(({ wf }) => reportJobs(wf).length > 0) | ||
|
|
||
| describe('unattended workflow failures are reported', () => { | ||
| it('discovers the known reporters', () => { | ||
| expect(reporters.map(({ path }) => path)).toEqual( | ||
| expect.arrayContaining(EXPECTED_REPORTERS), | ||
| ) | ||
| }) | ||
|
|
||
| it('every scheduled workflow calls the reporter or says why not', () => { | ||
| const silent = workflows | ||
| .filter(({ wf }) => triggers(wf).schedule) | ||
| .filter(({ wf }) => reportJobs(wf).length === 0) | ||
| .map(({ path }) => path) | ||
| .sort() | ||
| expect(silent).toEqual(Object.keys(REPORTS_ELSEWHERE).sort()) | ||
| }) | ||
|
|
||
| for (const { path, wf } of reporters) { | ||
| it(`${path} reports a failure in any of its jobs`, () => { | ||
| const jobs = reportJobs(wf) | ||
| expect(jobs).toHaveLength(1) | ||
| const [[name, job]] = jobs | ||
| const others = Object.keys(wf.jobs) | ||
| .filter((other) => other !== name) | ||
| .sort() | ||
| expect([job.needs].flat().sort()).toEqual(others) | ||
| expect(job.if).toBe(CONDITION) | ||
| expect(job.permissions).toEqual({ issues: 'write' }) | ||
| expect(String(job.with?.guidance ?? '').trim()).not.toBe('') | ||
| }) | ||
| } | ||
|
|
||
| it('the reporter keeps one open issue per workflow and branch', () => { | ||
| const wf = readWorkflow(REPORTER.slice(2)) | ||
| expect(triggers(wf).workflow_call?.inputs?.guidance?.required).toBe(true) | ||
| const steps = Object.values(wf.jobs).flatMap((job) => job?.steps ?? []) | ||
| const run = steps.map((step) => String(step?.run ?? '')).join('\n') | ||
| expect(run).toContain('gh issue create') | ||
| expect(run).toContain('gh issue comment') | ||
| // The issue is found by its title, so the title must name the caller. | ||
| const env = Object.assign({}, ...steps.map((step) => step?.env ?? {})) | ||
| expect(env.TITLE).toContain(expr('github.workflow')) | ||
| expect(env.TITLE).toContain(expr('github.ref_name')) | ||
| }) | ||
| }) |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Filter the issue lookup by the
needs-triagelabel and pass--limit.gh issue listreturns 30 issues by default, and--searchranks by relevance. If many open issues match the title words, the exact-title issue can fall outside that window. The workflow then opens a duplicate. This weakens the "one issue per workflow and branch" contract. Add--limit 100. Also,--search "in:title ..."is a word match, so the exactjqfilter is the right guard.Proposed fix
📝 Committable suggestion
🤖 Prompt for AI Agents