Conversation
A multi-file change asked the whole-change rubric once over a concatenated diff, so the correctness score came back flat and the gate escalated even when the per-file facts were easy to judge (jkudish#42). Both tools now take files: [{ path, diff }] instead of diff (exactly one of the two). The tool still writes the rubric questions itself; they are asked once per file, all in the same request, each scoped to its own files[i] and path. Action composes in code: the change is auto only when every file is auto (worstAction), composite is the file mean, safe_to_apply the file minimum, and limiting names the file and rubrics that bound the decision. Per-file results are returned in full; nothing is averaged into invisibility. Bounds: 16 files and 200,000 characters in aggregate, rejected before any model call; per-file diffs truncate at the same 50,000-character cap as diff, and truncated input still never returns auto. Claim questions in per-file gates name the file diffs instead of the diff. Whole-change diff behavior is unchanged.
Owner
|
Thank you for this. Merged with three amendments from review:
Thanks again! |
Owner
|
Merged onto main (rebased; the amendments above are part of the merged commits). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #42.
What
Both
jev_reviewandjev_gatenow takefiles: [{ path, diff }]instead ofdiff(exactly one of the two, enforced with a fixed-string error). The tool still writes the rubric questions itself — they are asked once per file, all in the same request, each scoped with aJudge only files[i] ("path"), not the change as a wholesentence, so the vague-question failure mode from the TypeSafe docs stays fixed.The action composes in code, per the issue's proposal:
autoonly when every file isauto(worstAction);compositeis the file mean,safe_to_applythe file minimum;limitingnames the file and rubrics that bound the decision (building on the Expose review score distributions and the rubric limiting automatic acceptance #20/Preserve review score distributions; reason codes for every action path #26 limiting-rubrics vocabulary);filesarray — nothing is averaged into invisibility, a weak file stays visible as itself.Per-file projection reuses
projectReviewHalfunchanged (slice thefile_i_*answers back to bare rubric keys), so validation, thresholds, reason codes, and limiting rubrics behave exactly as in whole-change mode.jev_gate's claim questions name the file diffs instead of the diff in per-file mode; claims and verification are untouched.Bounds mirror the gate's evidence discipline: up to 16 files, 200,000 characters in aggregate rejected before any model call; per-file diffs truncate at the same 50,000-character cap as
diff, and truncated input still never returnsauto.Testing
npm run typecheck,npm run build,npm test: 245 tests (all 239 pre-existing pass unchanged — whole-change behavior is untouched). New coverage: one request with per-file keys only and the scoped instructions; auto composition; the weak file naminglimitingwith file and rubric; exactly-one-ofdiff/files; the aggregate budget rejecting without a request; per-file truncation demoting auto; and the gate composing per-file review with claims in one call.Docs
README: bullet in
jev_review(with the mode semantics) and one injev_gate.