LLM-powered pull request review as a GitHub Action. One consolidated review per PR with inline findings, powered by any model on OpenRouter.
- Reviews the PR diff and traces changed code into its callers — regressions and pre-existing bugs in affected code are findings, not noise
- Reads your repo's conventions file (
AGENTS.mdby default) and reviews against it - Posts one consolidated PR review when there are inline findings, with comments anchored to diff lines; completed runs also update a status comment, including when there are no findings
- Structured output end to end: every finding carries a category, severity, confidence, and a concrete failure scenario
- Drops non-findings before they post — findings that conclude "no bug here" (an
N/A — …title, a "no action needed" suggestion, a "…is correct" title) are filtered deterministically - Model-agnostic via OpenRouter — pick your model, see your per-call costs; every finding comment carries an
umm-actually · <model>byline naming the model that produced that finding - Findings outside the diff on files the model was given (e.g. callers of changed code), findings in a changed file whose diff path can't take an inline comment, and every in-diff finding when GitHub rejects the inline review are posted as standalone comments on the PR
- PRs with oversized diffs are skipped gracefully with a body-only review stating the reason
- Reports as its own branded check run in the PR checks list — the App's avatar, the outcome as the check title (findings count, clean pass, or skip reason), and a details page carrying the summary and per-run cost
- Surfaces the review context in the workflow job summary — files seen, how much of the conventions file reached the model, which
priority_docswere included, the token budget breakdown, and how many findings each filter step dropped
umm-actually runs as a Docker-based action. It needs a GitHub token (for fetching the diff and posting the review) and an OpenRouter API key.
For the best experience, use a GitHub App installation token so reviews are attributed to a bot identity rather than a personal account.
Granting the App Checks: Read & write additionally puts the review in the PR checks list as a branded check run (the App's avatar instead of the generic Actions logo). The permission is optional — without checks: write on the token, the review runs unbranded and everything else works the same. The usage example below requests no permission narrowing on the token step, so the token picks up the Checks scope automatically once the App grants it; a workflow that does narrow permissions must list checks: write explicitly (and only once the App has the grant — requesting an ungranted permission fails the token step).
name: Review
on:
pull_request:
types: [opened, synchronize, reopened, ready_for_review]
issue_comment:
types: [created]
permissions:
contents: read
concurrency:
group: review-${{ github.event.pull_request.number || github.event.issue.number }}
cancel-in-progress: true
jobs:
review:
runs-on: ubuntu-latest
if: >-
github.event_name == 'pull_request' ||
(
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
contains(github.event.comment.body, '@umm review')
)
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
id: app-token
with:
client-id: ${{ secrets.UMM_CLIENT_ID }}
private-key: ${{ secrets.UMM_PRIVATE_KEY }}
- uses: aliasunder/umm-actually@v0
with:
github_token: ${{ steps.app-token.outputs.token }}
openrouter_api_key: ${{ secrets.OPENROUTER_KEY }}The @umm review comment trigger lets you re-request a review on any PR by commenting. The issue_comment event fires for PR comments — the if condition filters to PRs only.
An empty value, such as an unset repo variable, selects the input's default, so a workflow can wire repo variables without fallbacks of its own. There are two exceptions: github_token and openrouter_api_key are required, and an empty priority_docs disables priority docs.
| Input | Default | Description |
|---|---|---|
github_token |
(required) | Token for fetching the diff and posting the review. A GitHub App installation token keeps the bot identity. |
openrouter_api_key |
(required) | OpenRouter API key |
model |
anthropic/claude-sonnet-4-6 |
OpenRouter model slug exactly as listed on openrouter.ai/models |
fallback_model |
"" |
Model tried after the primary model's attempts fail (up to two per model). A timed-out primary attempt moves to this model immediately, without retrying the primary. Auth and credit errors (HTTP 401/402/403) stop the review without trying it. Empty = no fallback |
request_timeout_seconds |
900 |
Deadline for one model request, in seconds, capped by the time left in review_timeout_seconds. When it elapses, the attempt is recorded as timeout and the review moves to fallback_model if one is set; a timeout on the last model (the primary, when no fallback is set) may be retried once. The HTTP call is aborted best-effort: a request the provider keeps serving is still billed, and its cost-summary row shows n/a. Empty = 900 |
review_timeout_seconds |
1500 |
Shared model-work budget from action start, in seconds, including context preparation, all phases, retries, fallbacks and cost lookups. Completed phases post when it expires, with incomplete coverage identified; the run fails if none completed. Leave room below the workflow job timeout for setup and GitHub publication. External cancellation stops the run at once; unlike the deadline, it does not post completed phases. Empty = 1500 |
max_findings |
"" (uncapped) |
Cap on posted findings, highest severity first. Empty = all validated findings post. |
severity_threshold |
low |
Minimum severity to post: low | medium | high | critical |
conventions_file |
AGENTS.md |
Repo-relative path to the conventions file included in the prompt (truncated at conventions_budget_tokens). When the PR also changes the file, its full text is sent at most once, never twice |
conventions_budget_tokens |
8000 |
Approximate token budget for the conventions file section of the prompt. A file over the budget is truncated to its beginning, with a notice in the prompt. To send the whole file, also list it in priority_docs, which reads it in full when the context budget allows. The job summary reports every truncation. The status comment and the check run's details page also report it when the whole file never reached the model. They also report it when the whole file reached the model only because this PR changes the file (or, for a JavaScript/TypeScript conventions file, a file it imports) and the file is not listed in priority_docs, since later PRs will see just the truncated beginning. A listed file that stops fitting is reported on the PR where it does not fit. Empty = the 8000 default |
phases |
combined |
How the review dimensions are split across model calls: combined (one call carrying every dimension), parallel (three focused calls at once — each reads deeper, wall-clock time is the slowest call, and prompt tokens roughly triple), or sequential (the same three calls in order, each seeing the earlier calls' findings). When two calls report findings on overlapping lines of the same file, only the higher-severity finding is kept. Empty = combined |
context_budget_tokens |
300000 |
Approximate token budget for the diff plus every file sent in the prompt. A diff over half the budget skips the review. The conventions file has its own budget, conventions_budget_tokens |
trace_related_files |
true |
Add files beyond the diff to the prompt: JavaScript/TypeScript files that import a changed file by relative path (to catch caller regressions) and .md/.json files that mention a changed file (to catch stale docs). Does not affect priority_docs |
priority_docs |
README.md |
Comma-separated repo-relative paths sent to the model in full, with 10% of context_budget_tokens reserved for them, independent of trace_related_files. Docs not changed in the PR are read first, then changed ones, each in listed order, before changed files can use the budget. Reserved space they don't use goes back to changed files. A doc that doesn't fit the reserve is retried at the end with any budget left, and left out if it still doesn't fit. A doc already sent in full by another part of the prompt is not read again; a changed doc the prompt shows only as a diff can still be read here in full. A changed file excluded from the review diff stays excluded — diff_exclude_paths and linguist rules win over this list. Empty = disabled |
max_scan_files |
5000 |
Maximum files each repo scan collects before stopping. There are two scans, both run only when trace_related_files is on: JavaScript/TypeScript files for import tracing and .md/.json files for doc mentions. Files of other types or over max_scan_bytes don't count |
max_scan_bytes |
524288 |
Maximum byte size of a single file in the repo scans; larger files are skipped. It also caps the root .gitattributes; a larger one is ignored, so none of its linguist-generated rules apply |
max_related_files |
15 |
Maximum import-traced related files to include in review context |
max_related_docs |
10 |
Maximum mention-matched documentation files to include in review context. priority_docs entries are never mention-matched and don't count toward this cap |
exclude_paths |
"" |
Comma-separated folder paths, relative to the repo root, excluded from the repo scans (see max_scan_files). Files under these folders are invisible to import-tracing and doc-mention matching. Entries match folders only, by full path: fixtures skips the top-level fixtures/ but not src/fixtures/. The scans also always skip folders whose names start with . and any folder named node_modules, dist, build, out, or coverage. Changed files in the PR diff and priority_docs are never excluded. Empty = no exclusions. Example: evals, fixtures, test/fixtures |
diff_exclude_paths |
"" (built-in list) |
Comma-separated folder prefixes or globs removed from the review diff before the token budget check. Excluded files are listed by name in the review output, but their content is not reviewed — excluded is not vetted. Supplied patterns extend a built-in default list of generated artifacts (ecosystem lockfiles, *.min.js, *.min.css, *.map — the classes GitHub's linguist auto-collapses; snapshots are deliberately not defaulted, since GitHub renders them expanded). A leading none drops the defaults: none alone disables the built-in list (.gitattributes linguist-generated exclusions are governed separately by respect_linguist_generated), none, evals/** replaces the list. Unlike exclude_paths, empty means the default list, not "no exclusions". Patterns with more than 2 * in one path segment are rejected (** segments exempt) — glob matching backtracks exponentially on such shapes |
respect_linguist_generated |
true |
Also exclude changed files the repo's root .gitattributes marks linguist-generated=true (nested .gitattributes files are not read). Negated entries (-linguist-generated) keep a file reviewable even when the default diff_exclude_paths list matches it; explicitly supplied diff_exclude_paths patterns always win. Rules are read from the PR head, so a PR changing .gitattributes reviews under its own rules — every exclusion is named in the review output |
cost_summary |
true |
Write a per-run cost report (model, prompt/completion tokens, USD) to the workflow job summary and the check run's details page |
pr_number |
"" |
PR number override — required only when the triggering event does not identify a PR directly. Empty = taken from the event |
| Output | Description |
|---|---|
findings_count |
Number of new findings posted this run. Non-findings, findings on files the model never saw or that diff exclusion removed, duplicates, findings already posted on an earlier run, and findings below severity_threshold are removed before max_findings applies; a finding whose post failed is not counted |
review_url |
URL of the submitted review; empty when no review was posted |
model_used |
Model(s) that answered, as named in OpenRouter's response for each completed phase (the routed model, so a phase that fell back reports the fallback model), comma-separated when phases used different models. Empty when the review was skipped |
skipped_reason |
Non-empty when the review was skipped (e.g. diff too large) |
- Resolves the PR from the triggering event (supports
pull_request,pull_request_target, andissue_commentevents); once the PR context is known, opens a check run under the token's identity (best-effort — skipped when the token lackschecks: write) - Fetches the unified diff via the GitHub API — PRs that exceed the API's diff size limit are skipped
- Reads the conventions file, priority docs within their reserved budget, and changed files; then traces imports and scans docs (
.md,.json) with the remaining budget - Builds a structured prompt that wraps PR content in tags with a random per-prompt suffix, so PR text can't forge a closing tag to place text outside its tag; on re-runs, prior bot comment bodies are included so the model can self-suppress conceptual duplicates. Sends one request per review phase to OpenRouter (see the
phasesinput for dispatch modes) - Validates each response against a strict Zod schema. Each model gets up to two attempts, then
fallback_model, if set, is tried (see that input for timeouts and auth errors). A phase whose attempts all fail, or that reaches the shared review deadline, is named on the status comment and the check run while the other phases' findings still post; the run fails only when no phase completes - Drops non-findings (see Non-finding filter) and findings on files the model was never given or that diff exclusion removed (see Unknown-file filter), collapses findings that two phases reported on overlapping lines of the same file, then on re-runs deduplicates against previously posted bot comments (three-tier: positional match by hidden HTML anchor, content match by title similarity within 50 lines in the same file, or title-only match by high title similarity across any file)
- Filters remaining findings by severity threshold, drops the less severe of two same-category findings that overlap in the same file (on a tie, the one on the later line), and caps if configured
- Maps findings to inline PR review comments anchored to diff lines, with a snap-to-nearest-hunk fallback
- Posts one review with inline comments (its body is only a hidden marker); beyond-diff findings, findings in a changed file whose diff path can't take an inline comment, and every in-diff finding when GitHub rejects the inline review post as standalone PR comments; every run upserts a status comment with cross-run totals
- Completes the check run with the outcome — the conclusion grades the run, not the code:
successfor any completed review (with or without findings — the count is in the check title, and a review that lost a phase says so there too)neutralfor a skipfailureonly when the pipeline itself errorscancelledwhen the workflow run is cancelled mid-review (the action closes its own check on the way out instead of leaving it in progress)
Models sometimes report "findings" that conclude the code is fine — titled N/A — … or …is correct, with suggestions like "No action needed". The system prompt prohibits these, but models don't always comply, so every finding also passes a deterministic filter before cross-phase and cross-run deduplication, the severity threshold, and the cap. A finding is dropped when:
- its title, failure_scenario, or suggestion starts with a non-finding signal —
N/A,not applicable,placeholder, or a separator-delimited confirmation phrase (none — …,not a finding — …,no failure — …,no concrete failure scenario — …,no bug — …,no (further) action needed — …,no change needed — …) - its title ends with a declarative confirmation —
…is corrector…is accurate— or starts with a prior-finding resolution confirmation —Prior (bot) finding(s) addressed/resolved/fixed … - its title contains a self-negating phrase followed by a separator or end-of-title —
…not actionable,…not a defect,…not a(n) (real) issue/bug/problem - its failure_scenario ends with a leaked conclusion —
…no bug/…no bug here,…not actionable/…not a defect/…not a(n) (real) issue/bug/problem, or…analysis was wrong
The patterns are deliberately anchored (start-of-field, end-of-field, or separator-delimited) so real findings survive: a scenario like "No failure occurs until the third retry…" or "None of the guards catch this input" never matches. Dropped counts are logged per run (per-phase finding filters applied to model output) and appear in the job summary.
A finding's file (the path it is filed on) must name a file the model was given. A finding on any other path is ungrounded, because the model saw nothing there, and is dropped before cross-phase and cross-run deduplication, the severity threshold, and the cap. The model was given:
- paths in diff headers, including renamed-from and deleted paths (findings on renamed-from and deleted paths post as standalone comments)
- changed files, import-traced related files, mention-matched docs, and priority docs
- the conventions file
Files removed by diff_exclude_paths or linguist rules are listed by path at the end of the diff, but their content is withheld, so findings on them are dropped too and reported as excluded-file drops. When the conventions file is excluded, it is still sent as the review's conventions, but findings on it are dropped like those on any other excluded file. The filter compares paths and reports drops as follows:
- Paths are normalized before comparison: surrounding whitespace,
.segments,..segments that stay inside the repository, repeated slashes, and a leading or trailing/don't affect the match (./src/x.ts,/src/x.ts, andsrc/x.tsmatch). Matching is case-sensitive, and only whole paths match: a bare filename or a directory prefix does not. - A path containing
"reaches the model with each"written as"in the tags that wrap each file's content and the conventions file. A diff header prints the decoded path, so each"appears as written, and the filter matches that spelling as written. A path whose decoding was rejected keeps the diff's escaped spelling in its header. A model that copies the path from a tag writes"intofile, so afilethat matches no given path as written, but matches once each"becomes", is kept and posts under that decoded path. Each such rewrite is logged at debug level (resolved escaped finding file to a prompt path) with the model's spelling and the decoded path. - Each drop is logged as a warning with the file, line, and category:
dropping finding: file not in prompt contextfor a file the model was never given, anddropping finding: file excluded from the review difffor an excluded file. The job summary counts each kind per run, in theDropped as unknown fileandDropped as excluded filerows.
To see the model's reasoning behind its findings, set the LOG_LEVEL environment variable (not an action input) on the umm-actually step:
- uses: aliasunder/umm-actually@v0
env:
LOG_LEVEL: debug
with:
github_token: ${{ steps.app-token.outputs.token }}
openrouter_api_key: ${{ secrets.OPENROUTER_KEY }}- What it logs: the
analysisfield of each completed review phase, as one JSON line with the messagereview phase analysisin the umm-actually step's log. The defaultphases: combinedruns one phase per review;parallelandsequentialeach run three. - Matching a finding: the model is asked to write one analysis line per changed file, plus one line per finding outside the diff in the form
<finding title> — <path>: "<quoted passage>"(for missing content, the nearest heading followed by(missing here)). Search the analysis for an outside-the-diff finding's title, or for an inline finding's file path. - Findings already posted: the analysis is logged only on runs with debug on. Turn it on and trigger a new review with a push or an
@umm reviewcomment (the Actions Re-run button replays the original run's workflow file, without the newenv:line); the new review is a new model call and may not report the same finding. An@umm reviewcomment runs the workflow file from the default branch, so theenv:line must be merged there first. - Who can read it: the analysis quotes repository content, and anyone who can read the workflow's logs can read it. On a public repository, that is everyone.
- Values:
debug,info(the default),warn, orerror, in any letter case. Any other value falls back toinfo.
- Under consideration: a
read_filetool, so the model can read more files before finalizing findings - Later: bounded multi-step investigation, where the model uses tools to explore the repo before reporting
See the CHANGELOG for what each release shipped.