Summary
jev_review and the review half of jev_gate ask their rubric questions once over the whole diff. On a change that spans several files, the correctness Score comes back flat and the gate escalates, even when the per-file facts are easy to judge. Would you consider a per-file (or per-hunk) review mode, where the tool still writes the rubric questions itself but asks them once per file, all in the same request, and composes the action in code?
What we saw
A jev_gate call (v0.10.1, jev-1.13.0) on one change spanning five files returned:
"correctness": { "score": 0.68, "confidence": 0, "probabilities": { "0": 0.48, "1": 0.37, "2": 0.15 } },
"safe_to_apply": 0.11,
"action": "escalate",
"limiting_rubrics": ["correctness"]
The claim half of the same call went well: all 5 claims came back verified, 3 of them for automatic acceptance. When we then checked the same rules as 20 narrow claims, each matched to its own evidence item (7 items, our test logs), in one jev_verify call, 19 of 20 came back verified for automatic acceptance.
Part of this was our fault: we sent a multi-file change to a tool whose question is about "this change" as a whole. But the TypeSafe docs describe exactly this failure mode: "vague questions give mushy, uncalibrated scores" (cookbooks/sde_cascade), and the recommended shape is "Break broad judgments into narrow, typed questions with explicit instructions and criteria" (concepts/how-to-build-with-system-one).
Proposal
Keep the rubric as it is, and add an optional input to jev_review and jev_gate:
This keeps your line from #31: the tool still constructs and validates its own questions; the caller only splits the change.
Why it matters to us
We route agents to jev_review/jev_gate before a commit. Today we have to tell them "only for one small change in one file; otherwise use jev_verify with narrow claims". A per-file mode would let the review tools cover real multi-file commits.
Happy to test a branch against our labelled cases.
Summary
jev_reviewand the review half ofjev_gateask their rubric questions once over the wholediff. On a change that spans several files, the correctness Score comes back flat and the gate escalates, even when the per-file facts are easy to judge. Would you consider a per-file (or per-hunk) review mode, where the tool still writes the rubric questions itself but asks them once per file, all in the same request, and composes the action in code?What we saw
A
jev_gatecall (v0.10.1,jev-1.13.0) on one change spanning five files returned:The claim half of the same call went well: all 5 claims came back verified, 3 of them for automatic acceptance. When we then checked the same rules as 20 narrow claims, each matched to its own evidence item (7 items, our test logs), in one
jev_verifycall, 19 of 20 came back verified for automatic acceptance.Part of this was our fault: we sent a multi-file change to a tool whose question is about "this change" as a whole. But the TypeSafe docs describe exactly this failure mode: "vague questions give mushy, uncalibrated scores" (
cookbooks/sde_cascade), and the recommended shape is "Break broad judgments into narrow, typed questions with explicit instructions and criteria" (concepts/how-to-build-with-system-one).Proposal
Keep the rubric as it is, and add an optional input to
jev_reviewandjev_gate:files: an array of{ path, diff }items (instead of, or next to, onediffstring).primitives).autoonly when every file isauto, and the result names the limiting file and rubric (building on Expose review score distributions and the rubric limiting automatic acceptance #20 / Preserve review score distributions; reason codes for every action path #26).This keeps your line from #31: the tool still constructs and validates its own questions; the caller only splits the change.
Why it matters to us
We route agents to
jev_review/jev_gatebefore a commit. Today we have to tell them "only for one small change in one file; otherwise usejev_verifywith narrow claims". A per-file mode would let the review tools cover real multi-file commits.Happy to test a branch against our labelled cases.