Skip to content

feat: filter non-findings and constrain comment verbosity - #9

Merged
aliasunder merged 5 commits into
mainfrom
worktree-feat-noise-filter
Jul 13, 2026
Merged

aliasunder merged 5 commits into
mainfrom
worktree-feat-noise-filter

Conversation

@aliasunder

@aliasunder aliasunder commented Jul 13, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Prompt hardening — new OUTPUT_DISCIPLINE section in the system prompt constrains field content (imperative titles <80 chars, 1-3 sentence descriptions without code traces, concrete failure scenarios) and prohibits "no bug" conclusions from appearing as findings
  • Post-LLM filter — deterministic regex filter on failure_scenario catches non-findings the prompt didn't prevent (N/A, None, "this is correct", etc.), integrated as Step 12.5 in orchestrate between selection and posting
  • Collapsible suggestions — diff blocks wrapped in <details> to keep the main comment body focused on the problem and failure scenario

Motivated by PR #8 where ~9 of 14 inline comments were non-findings that concluded "this is correct / no bug / N/A" and real findings contained 200-word code-path traces instead of concise problem statements.

Test plan

  • 18 new tests for filterNonFindings covering all prefix patterns (N/A, None, Not applicable, Placeholder), body patterns (no actual bug, this is correct, working as designed, by design, correct behavior, no bug found), and edge cases (mid-sentence "None", empty array, mixed sets)
  • 1 new test for OUTPUT DISCIPLINE ordering in system prompt
  • 1 new orchestrate integration test with full deterministic assertions (mixed non-finding + real finding fixture → verify complete result object and submitReview call params)
  • 3 updated comment-mapping tests for <details> wrapper
  • 231 total tests pass, build clean, lint clean

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Review results now exclude responses that indicate no actual issue, preventing false-positive findings.
    • Reviews with only non-findings correctly follow the zero-findings behavior.
  • Improvements

    • Suggested fixes now appear in collapsible “Suggested fix” sections for clearer review comments.
    • Output guidance better ensures findings describe concrete problems and include the required details.

Reduce review noise via three changes:

1. Prompt hardening — OUTPUT DISCIPLINE section constraining field content
   (imperative titles, 1-3 sentence descriptions, concrete failure scenarios)
   and prohibiting "no bug" conclusions from appearing as findings.

2. Post-LLM filter — deterministic regex filter on failure_scenario catches
   non-findings the prompt didn't prevent (N/A, None, "this is correct", etc).
   Integrated as Step 12.5 in orchestrate between selection and posting.

3. Collapsible suggestions — diff blocks wrapped in <details> to keep the
   main comment body focused on the problem and failure scenario.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @aliasunder, you have reached your weekly rate limit of 2500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@umm-actually umm-actually Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed — 1 finding(s) posted as inline comments.


umm-actually · deepseek/deepseek-v4-pro-20260423

Comment thread src/review/filter-non-findings.ts Outdated
@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@aliasunder, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 45 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: b64ffc70-0d39-4637-9955-75a0c2716a71

📥 Commits

Reviewing files that changed from the base of the PR and between 8b46808 and 7d2fdfa.

📒 Files selected for processing (7)
  • src/__tests__/orchestrate.test.ts
  • src/orchestrate.ts
  • src/review/__tests__/comment-mapping.test.ts
  • src/review/__tests__/filter-non-findings.test.ts
  • src/review/__tests__/prompt.test.ts
  • src/review/comment-mapping.ts
  • src/review/filter-non-findings.ts
📝 Walkthrough

Walkthrough

The review pipeline now constrains and filters non-findings before submission. Orchestration reports only retained findings, suggestion diffs render in collapsible sections, and tests cover filtering, prompt ordering, orchestration, and fence sizing.

Changes

Finding quality and review output

Layer / File(s) Summary
Non-finding detection and output discipline
src/review/filter-non-findings.ts, src/review/prompt.ts, src/review/__tests__/*
Adds prompt rules and regex-based filtering for non-findings, with coverage for matching, passthrough, edge cases, and prompt section ordering.
Filtered orchestration submission
src/orchestrate.ts, src/__tests__/orchestrate.test.ts
Filters threshold-selected findings before constructing review output and verifies submission uses only retained findings.
Collapsible suggestion rendering
src/review/comment-mapping.ts, src/review/__tests__/comment-mapping.test.ts
Wraps suggested diffs in a <details> block and validates fence sizing when suggestions contain backticks.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: non-finding filtering and reduced comment verbosity.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch worktree-feat-noise-filter

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (2)
src/__tests__/orchestrate.test.ts (1)

657-723: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add an assertion on the new "filtered non-findings" log entry.

The test exercises the droppedAsNonFinding > 0 branch in orchestrate.ts (Step 12.5) but never asserts on logger.messages, so a regression that drops or malforms that log call wouldn't be caught here.

✅ Proposed addition
       expect(stubs.submitReviewCalls).toHaveLength(1)
       expect(first(stubs.submitReviewCalls)).toEqual({
         prNumber: fixturePrContext.prNumber,
         commitId: fixturePrContext.headSha,
         body: mixedBody,
         comments: mixedMapped.comments,
         fallbackBody: mixedFallbackBody,
       })
+      expect(logger.messages).toContainEqual(
+        expect.objectContaining({
+          level: "info",
+          message: "filtered non-findings",
+          data: expect.objectContaining({ droppedAsNonFinding: 1 }),
+        }),
+      )
     })
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/__tests__/orchestrate.test.ts` around lines 657 - 723, Add an assertion
in the “filters non-findings from LLM output” test that checks logger.messages
contains the Step 12.5 filtered non-findings log entry, including the expected
droppedAsNonFinding count and message structure. Keep the existing orchestration
and submission assertions unchanged.
src/review/filter-non-findings.ts (1)

8-16: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Heuristic can silently drop legitimate findings, with no per-item visibility.

NON_FINDING_BODY will match a genuine bug description that legitimately contains phrases like "correct behavior" or "this is correct" while contrasting it with the actual (buggy) behavior — e.g. "the correct behavior is to reject empty strings, but this is correct only for the happy path; empty input bypasses the check." Such a finding would be silently dropped in filterNonFindings, and orchestrate.ts only logs an aggregate droppedAsNonFinding count (Line 259-261), not which findings were removed, making false positives undetectable in production.

Consider logging the dropped finding's file/line/title (not full description) alongside the count for debuggability.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/review/filter-non-findings.ts` around lines 8 - 16, The broad
NON_FINDING_BODY heuristic in isNonFinding can discard legitimate findings, and
dropped items are not individually visible. Narrow the matching logic to avoid
treating contextual phrases such as “correct behavior” or “this is correct” as
sufficient on their own, and update filterNonFindings/orchestrate logging to
record each dropped finding’s file, line, and title while retaining the
aggregate droppedAsNonFinding count.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/review/__tests__/comment-mapping.test.ts`:
- Around line 285-287: Update the assertion in the comment-mapping test to
validate the complete deterministic rendered body, replacing toContain with an
exact equality check such as toBe (or isolate the specific finding before
comparing). Preserve the expected suggestion markup while ensuring unexpected
surrounding content cannot pass the test.

In `@src/review/__tests__/filter-non-findings.test.ts`:
- Around line 28-36: Update the fixture in the test for the no-slash NA prefix
branch of filterNonFindings so its failure_scenario does not contain wording
matched by NON_FINDING_BODY. Use text that only matches NON_FINDING_PREFIX,
while preserving the expected dropped finding count and empty findings result.

In `@src/review/filter-non-findings.ts`:
- Around line 8-11: Add documentation comments directly above NON_FINDING_PREFIX
and NON_FINDING_BODY explaining each regex’s matching intent, including prefix
word-boundary anchoring and the covered non-finding phrases. Do not change the
regex patterns or surrounding filtering behavior.

---

Nitpick comments:
In `@src/__tests__/orchestrate.test.ts`:
- Around line 657-723: Add an assertion in the “filters non-findings from LLM
output” test that checks logger.messages contains the Step 12.5 filtered
non-findings log entry, including the expected droppedAsNonFinding count and
message structure. Keep the existing orchestration and submission assertions
unchanged.

In `@src/review/filter-non-findings.ts`:
- Around line 8-16: The broad NON_FINDING_BODY heuristic in isNonFinding can
discard legitimate findings, and dropped items are not individually visible.
Narrow the matching logic to avoid treating contextual phrases such as “correct
behavior” or “this is correct” as sufficient on their own, and update
filterNonFindings/orchestrate logging to record each dropped finding’s file,
line, and title while retaining the aggregate droppedAsNonFinding count.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: ae4d37ad-82c8-4933-8052-57377ac473c8

📥 Commits

Reviewing files that changed from the base of the PR and between 5342f4f and 8b46808.

📒 Files selected for processing (8)
  • src/__tests__/orchestrate.test.ts
  • src/orchestrate.ts
  • src/review/__tests__/comment-mapping.test.ts
  • src/review/__tests__/filter-non-findings.test.ts
  • src/review/__tests__/prompt.test.ts
  • src/review/comment-mapping.ts
  • src/review/filter-non-findings.ts
  • src/review/prompt.ts

Comment thread src/review/__tests__/comment-mapping.test.ts Outdated
Comment thread src/review/__tests__/filter-non-findings.test.ts
Comment thread src/review/filter-non-findings.ts Outdated

@aliasunder aliasunder left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Phase 1 review (correctness + security/performance) for "feat: filter non-findings and constrain comment verbosity."

Reviewed all 8 changed files in full, plus src/orchestrate.ts, src/review/prompt.ts, src/review/comment-mapping.ts, src/review/finding.ts, src/review/select-findings.ts, and the surrounding test files for integration context.

3 findings — all relate to the new filter-non-findings.ts and its integration. The <details> wrapper in comment-mapping.ts and the OUTPUT_DISCIPLINE prompt section are clean.

No ReDoS risk — both regexes are fixed-alternation, linear-time. No injection or data-leakage concerns.

The core concern: the NON_FINDING_BODY substring filter is unanchored and matches phrases that appear naturally in legitimate bug descriptions, not just in no-bug conclusions. This can silently drop real findings.


🔍 ship-check · PHASE 1: PR REVIEW · openrouter/z-ai/glm-5.2

Comment thread src/review/filter-non-findings.ts Outdated
Comment thread src/review/filter-non-findings.ts Outdated
Comment thread src/orchestrate.ts Outdated

@aliasunder aliasunder left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code quality pass — reviewed 4 changed non-test files (orchestrate.ts, filter-non-findings.ts, prompt.ts, comment-mapping.ts) against AGENTS.md conventions.

1 finding, HIGH confidence, in the Comments category.

Comments (1)

  • src/review/filter-non-findings.ts: The two regex constants NON_FINDING_PREFIX and NON_FINDING_BODY have no doc comments. AGENTS.md requires regex constants to carry doc comments explaining what they match. The filter's two-pattern strategy (prefix match vs. body match) is non-obvious from the regexes alone — a one-line comment per constant would make the intent readable without parsing the pattern.

Everything else passes: module layering (pure leaf in review/, named exports), naming (selectedByThreshold → selected rename chain in orchestrate.ts is clear), export style (named exports for a small-helper module), explicit return types on all exports, type over interface, and the <details> wrapper in comment-mapping.ts follows the existing long-template-literal style in that file.


🔍 ship-check · PHASE 2: CODE QUALITY · openrouter/z-ai/glm-5.2

Comment thread src/review/filter-non-findings.ts Outdated
Remove the unanchored body regex (NON_FINDING_BODY) and the "none" prefix
pattern — both produce false positives on real findings:

- "by design" matches "returns null by design, but the caller never
  null-checks" — a real bug with a qualifying clause
- "this is correct" matches "correct for ASCII but breaks on UTF-8"
- "None" matches "None of the guards catch this input"

The filter now only matches unambiguous non-finding prefixes: N/A, NA,
Not applicable, Placeholder. The prompt hardening steers models toward
these prefixes when there is no real finding.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@umm-actually umm-actually Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed — 1 finding(s) posted as inline comments.


umm-actually · deepseek/deepseek-v4-pro-20260423

Comment thread src/review/filter-non-findings.ts

@aliasunder aliasunder left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Phase 3: Test Audit

Audited 4 changed test files and ran coverage gap analysis on 4 changed production files.

Findings: 4 total (2 medium, 2 low)

Coverage gap analysis

  • filter-non-findings.ts — well covered overall: every regex alternative has a dedicated test, edge cases (empty, all-non-findings, all-real, mixed) are present, and a false-positive guard ("None" mid-sentence) confirms the prefix is start-anchored. One two-bar violation found (inline).
  • orchestrate.ts (Step 12.5) — filter wiring is tested end-to-end via the mixed-findings integration test. The hasFindings = false → buildZeroFindingsBody branch is already covered by the existing "all findings below threshold" test (line 582), so the filter-empties-everything path exercises the same branch. The droppedAsNonFinding logging branch is exercised but not asserted (inline).
  • prompt.ts (OUTPUT_DISCIPLINE) — section ordering is verified. Section headers are verified present. The key prohibitions (prefix list, body-phrase list) that form a contract with filter-non-findings.ts are not verified — only the headers (inline).
  • comment-mapping.ts (<details> wrapper) — wrapper structure (open tag, summary, fence, content, close tag) is verified via toContain. A full-body toEqual would additionally pin the wrapper's position, consistent with the sibling assertion at line 38 (inline, low).

By category

  • Two-bar violations: 1
  • Assertion quality: 2
  • Test hygiene: 0
  • Completeness: 1
  • Coverage regressions: 0

No findings were duplicated from Phase 1 (regex over-matching, "None" prefix, filter observability — design decisions) or Phase 2 (regex doc comments — code quality).


🔍 ship-check · PHASE 3: TEST AUDIT · openrouter/z-ai/glm-5.2

Comment thread src/review/__tests__/filter-non-findings.test.ts Outdated
Comment thread src/review/__tests__/prompt.test.ts
Comment thread src/__tests__/orchestrate.test.ts
Comment thread src/review/__tests__/comment-mapping.test.ts
Comment thread src/review/__tests__/filter-non-findings.test.ts Outdated
Comment thread src/review/__tests__/prompt.test.ts
Comment thread src/__tests__/orchestrate.test.ts
Comment thread src/review/__tests__/comment-mapping.test.ts
…ture clarity

- Add doc comment to NON_FINDING_PREFIX regex constant (AGENTS.md convention)
- Clean up NA test fixture text to avoid ambiguity with removed body patterns
- Assert logger.messages for the "filtered non-findings" log in orchestrate test
- Assert key OUTPUT_DISCIPLINE prohibitions in prompt test (not just headers)
- Upgrade comment-mapping suggestion test from toContain to full toBe assertion

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@aliasunder

Copy link
Copy Markdown
Owner Author

Bot review body findings — disposition

Sourcery (review PRR_kwDOTSnfAM8AAAABFwMR3w): Rate-limited, no findings.

CodeRabbit (review PRR_kwDOTSnfAM8AAAABFwM1OQ): Two nitpick findings — both duplicated as inline threads and handled there:

  1. Assert droppedAsNonFinding log → fixed in 792cb8c
  2. NON_FINDING_BODY false positives → body regex removed entirely in 7204628

ship-check Phase 1 (review PRR_kwDOTSnfAM8AAAABFwOrLg): 3 findings, all handled as inline threads:

  1. NON_FINDING_BODY unanchored matching → fixed (body regex removed)
  2. "none" prefix false positive → fixed ("none" removed from prefix)
  3. Filter observability → intentional design (log-only, not review body)

ship-check Phase 2 (review PRR_kwDOTSnfAM8AAAABFwPHbw): 1 finding (regex doc comments) → fixed in 792cb8c

ship-check Phase 3 (review PRR_kwDOTSnfAM8AAAABFwQNsw): 4 findings, all handled as inline threads:

  1. Two-bar violation in NA test → fixed (fixture cleaned up)
  2. Log assertion gap → fixed (logger assertion added)
  3. Prompt test coverage gap → fixed (key prohibitions asserted)
  4. toContain → toBe in comment-mapping → fixed (full body assertion)

umm-actually (reviews PRR_kwDOTSnfAM8AAAABFwMafg, PRR_kwDOTSnfAM8AAAABFwQAuw): Summary bodies only, findings are inline threads — all handled.


🔍 ship-check · pr-monitor · Opus 4.6 (1M context)

@aliasunder aliasunder left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Phase 4: Bug Check

Reviewed all 4 changed production files in full (orchestrate.ts, filter-non-findings.ts, prompt.ts, comment-mapping.ts) plus integration context (select-findings.ts, finding.ts).

4 findings — 1 medium-confidence (needs design decision), 3 low-confidence (flagged).

Summary

  1. Filter runs after cap — cap slots wasted on non-findings (orchestrate.ts:258) — medium confidence, needs design decision. If maxFindings=5 and 2 of the top 5 are non-findings, the filter drops them, leaving 3 posted. Real findings #6–7 were already dropped by the cap and could have filled those slots.

  2. Empty-string suggestion renders a confusing empty <details> block (comment-mapping.ts:41) — low/medium confidence, pre-existing gap worsened by this PR.

  3. Prompt prohibits "no bug" but filter doesn't match it (filter-non-findings.ts:11) — low confidence, description mismatch.

  4. Leading whitespace bypasses the anchored prefix regex (filter-non-findings.ts:8) — low confidence, input validation gap.

Verified all regex behavior with Node.js. CI checks job (Test step) passes — confirmed test file at HEAD matches implementation.


🔍 ship-check · PHASE 4: BUG CHECK · openrouter/z-ai/glm-5.2

Comment thread src/orchestrate.ts Outdated
Comment thread src/review/comment-mapping.ts
Comment thread src/review/filter-non-findings.ts
Comment thread src/review/filter-non-findings.ts
- Move filterNonFindings before selectFindings so cap slots aren't wasted
  on non-findings that would be dropped anyway
- Trim failure_scenario before regex test to handle leading whitespace
- Guard against empty-string suggestion rendering a confusing empty
  <details> block

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@aliasunder

Copy link
Copy Markdown
Owner Author

🔍 Phase 4 Bug Check — 3 findings (against current head 792cb8c)

The following findings from the bug-check phase did not land as inline comments due to an API formatting issue. Re-posting here as a consolidated comment.


1. Filter-after-cap ordering — non-findings consume maxFindings slots (medium)

File: src/orchestrate.ts:257-258

filterNonFindings (Step 12.5) runs after selectFindings (Step 12) applies the maxFindings cap. If the LLM returns non-findings ranked above real findings (LLMs may rate "this is correct" as high severity), selectFindings caps them in, then the filter drops them — leaving the user with fewer real findings than the cap allows.

Example: maxFindings=5, LLM returns 3 non-findings (severity: high) + 5 real findings (severity: medium). selectFindings keeps the 3 non-findings + 2 real findings (cap=5). filterNonFindings drops the 3, leaving only 2 real findings — 3 short of the cap.

Fix: Move filterNonFindings before selectFindings (swap Steps 12 and 12.5), or filter inside selectFindings before capping.

Category: Needs design decision — the PR description explicitly places the filter after selection. The fix is trivial (swap two blocks) but the policy question (before or after cap) needs a decision.


2. Empty suggestion renders empty <details> block (low/medium, pre-existing)

File: src/review/comment-mapping.ts:31

suggestionBlock guards against null (if (finding.suggestion === null) return "") but not empty string "". The Zod schema is suggestion: z.string().nullable() — empty string is a valid value. If the LLM returns suggestion: "", the function renders <details><summary>Suggested fix</summary>\n\n```diff\n\n```\n\n</details> — an empty collapsible diff block. The <details> wrapper makes this more visually prominent than the pre-PR bare empty code block.

Fix: if (finding.suggestion === null || finding.suggestion.trim() === "") return "" — or enforce .min(1) in the Zod schema.

Category: Pre-existing gap — the guard predates this PR; the <details> wrapper amplifies the visibility.


3. Leading whitespace bypasses ^-anchored prefix regex (low)

File: src/review/filter-non-findings.ts:8

NON_FINDING_PREFIX uses ^ (start-of-string anchor) without allowing leading whitespace. Zod does not trim failure_scenario, so " N/A — tests are valid." (leading space) bypasses the filter.

Fix: Change ^ to ^\s* in the regex, or .trim() the scenario before testing.

Category: Trivial fix — one-line change, no interface impact.


🔍 ship-check · PHASE 4: BUG CHECK · openrouter/z-ai/glm-5.2

@umm-actually umm-actually Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed — 1 finding(s) posted as inline comments.


umm-actually · deepseek/deepseek-v4-pro-20260423

Comment thread src/review/filter-non-findings.ts
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@umm-actually umm-actually Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed — 1 finding(s) posted as inline comments.


umm-actually · deepseek/deepseek-v4-pro-20260423

}

/** Matches failure_scenario prefixes that signal "no real finding" — start-anchored, case-insensitive, word-bounded. */
const NON_FINDING_PREFIX = /^(?:n\/?a|not applicable|placeholder)\b/i

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[medium/correctness] Add missing 'None' prefix to NON_FINDING_PREFIX regex (confidence: medium)

The OUTPUT_DISCIPLINE section of the system prompt lists "None" as one of the failure_scenario prefixes to avoid, and the PR description says the regex filter catches N/A, None, Not applicable, and Placeholder. The regex in filter-non-findings.ts only matches n/a (with optional slash), not applicable, and placeholder — it does NOT match None. A finding whose failure_scenario is just "None" (capitalized, without of trailing) will be classified as a real finding instead of being dropped.

Failure scenario: The LLM emits a finding with "failure_scenario": "None" (a single word matching what the prompt tells it to avoid). The regex /^(?:n\/?a|not applicable|placeholder)\b/i does not match "None" because None is not in the alternation. The finding passes through filterNonFindings and is selected by selectFindings, wasting a cap slot and appearing in the review.

Suggested fix
-const NON_FINDING_PREFIX = /^(?:n\/?a|not applicable|placeholder)\b/i
+const NON_FINDING_PREFIX = /^(?:n\/?a|none|not applicable|placeholder)\b/i

@aliasunder
aliasunder merged commit 889d97b into main Jul 13, 2026
9 checks passed
aliasunder added a commit that referenced this pull request Jul 22, 2026
* feat: extend non-finding filter to title + suggestion fields

Six confirmation findings escaped on PRs #10/#12 because the model routed
the "no real finding" signal into whichever field the filter didn't
inspect — the title ("N/A — …", "…is correct"), the suggestion
("No action needed — …"), or a rambling scenario's conclusion
("…No bug here."). The filter tested only failure_scenario prefixes.

Filter (deterministic layer):
- title, failure_scenario, and suggestion all tested against the existing
  prefix set plus a new separator-anchored confirmation set (none / no
  failure / no concrete failure scenario / no bug / no action needed /
  no change needed). The required separator ([—–:.-] or end-of-text)
  keeps real scenarios like "No failure occurs until…" and "None of
  the guards…" safe — and lets "none" return, resolving the PR #9
  prompt/filter drift in the filtering direction.
- Declarative confirmation titles ("…is correct" / "…is accurate")
  and leaked scenario conclusions ("…no bug here", "…analysis was
  wrong") drop via end-anchored suffixes.

Prompt (advisory layer): OUTPUT_DISCIPLINE bans declarative confirmation
titles + verification-task titles, and gains a suggestion-format line —
a suggestion of "no bug"/"no action needed"/"N/A" means no finding.

README: new Non-finding filter section documenting the drop contract and
the anchoring rationale; What-it-does + How-it-works + Shipped surfaces
updated in lockstep.

Accepted tradeoffs (user-approved): genuine "Placeholder text …" titles
and imperative "Ensure the timeout is correct" titles would drop —
zero observed instances in ~70 bot findings vs 6 observed escapes.

17 new tests on the real escaped strings (398 total); 3 mutation checks
verified (title-suffix rule, separator anchoring, title field removal
each fail exactly their covering tests).

* fix(review): document optional "here" in no-bug scenario suffix

The CONFIRMATION_SCENARIO_SUFFIX regex uses `\bno bug(?: here)?` which
matches both "...no bug" and "...no bug here", but the README only listed
the latter. Add both variants so the filter contract matches the code.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: use truthy check for optional reason attribute

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: cover CONFIRMATION_SCENARIO_SUFFIX 'analysis was wrong' and title-field CONFIRMATION_PREFIX

Two coverage gaps found during test audit:
- The 'analysis was wrong' alternative in CONFIRMATION_SCENARIO_SUFFIX had
  no end-of-string test (the existing test matched 'No bug here.' instead)
- CONFIRMATION_PREFIX on the title field was untested (only NON_FINDING_PREFIX
  was exercised on titles)

Both mutation-verified: removing each regex alternative fails exactly the
new test.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: replace prompt fragment assertions with full-constant checks

The OUTPUT_DISCIPLINE and ANCHORING_CONTRACT assertions used multiple
toContain calls with short phrases — violating the convention that
deterministic output gets exact assertions. Replaced with single
toContain of each full constant text.

AGENTS.md test conventions now explicitly address the large-string
case: assert the full section in one toContain, not multiple fragments.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@aliasunder
aliasunder deleted the worktree-feat-noise-filter branch July 25, 2026 18:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant