Skip to content

jev_review / jev_gate: optional per-file review mode - #47

Closed
Rex-Gao wants to merge 1 commit into
jkudish:mainfrom
Rex-Gao:per-file-review
Closed

Rex-Gao wants to merge 1 commit into
jkudish:mainfrom
Rex-Gao:per-file-review

Conversation

@Rex-Gao

@Rex-Gao Rex-Gao commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Implements #42.

What

Both jev_review and jev_gate now take files: [{ path, diff }] instead of diff (exactly one of the two, enforced with a fixed-string error). The tool still writes the rubric questions itself — they are asked once per file, all in the same request, each scoped with a Judge only files[i] ("path"), not the change as a whole sentence, so the vague-question failure mode from the TypeSafe docs stays fixed.

The action composes in code, per the issue's proposal:

Per-file projection reuses projectReviewHalf unchanged (slice the file_i_* answers back to bare rubric keys), so validation, thresholds, reason codes, and limiting rubrics behave exactly as in whole-change mode. jev_gate's claim questions name the file diffs instead of the diff in per-file mode; claims and verification are untouched.

Bounds mirror the gate's evidence discipline: up to 16 files, 200,000 characters in aggregate rejected before any model call; per-file diffs truncate at the same 50,000-character cap as diff, and truncated input still never returns auto.

Testing

npm run typecheck, npm run build, npm test: 245 tests (all 239 pre-existing pass unchanged — whole-change behavior is untouched). New coverage: one request with per-file keys only and the scoped instructions; auto composition; the weak file naming limiting with file and rubric; exactly-one-of diff/files; the aggregate budget rejecting without a request; per-file truncation demoting auto; and the gate composing per-file review with claims in one call.

Docs

README: bullet in jev_review (with the mode semantics) and one in jev_gate.

A multi-file change asked the whole-change rubric once over a concatenated
diff, so the correctness score came back flat and the gate escalated even
when the per-file facts were easy to judge (jkudish#42).

Both tools now take files: [{ path, diff }] instead of diff (exactly one
of the two). The tool still writes the rubric questions itself; they are
asked once per file, all in the same request, each scoped to its own
files[i] and path. Action composes in code: the change is auto only when
every file is auto (worstAction), composite is the file mean,
safe_to_apply the file minimum, and limiting names the file and rubrics
that bound the decision. Per-file results are returned in full; nothing
is averaged into invisibility.

Bounds: 16 files and 200,000 characters in aggregate, rejected before
any model call; per-file diffs truncate at the same 50,000-character cap
as diff, and truncated input still never returns auto. Claim questions
in per-file gates name the file diffs instead of the diff. Whole-change
diff behavior is unchanged.

jkudish commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Thank you for this. Merged with three amendments from review:

  • questions scope by index only, so file paths can't carry directives into the rubric
  • one combined budget (request + tests + all diffs), so a single call stays inside state limits
  • truncation is per-file: an intact file is never demoted for a truncated sibling

Thanks again!

jkudish commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Merged onto main (rebased; the amendments above are part of the merged commits).

@jkudish jkudish closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants