Skip to content

Add jev_audit: verify extracted values against their source - #48

Merged
jkudish merged 1 commit into
jkudish:mainfrom
Rex-Gao:audit-tool
Sep 28, 2026
Merged

jkudish merged 1 commit into
jkudish:mainfrom
Rex-Gao:audit-tool

Conversation

@Rex-Gao

@Rex-Gao Rex-Gao commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Implements #45 — the twelfth tool, proposed there with the research notes.

What

jev_audit audits extracted values against the text they claim to come from, before the values are trusted. One request per call carries the SDE cascade cookbook's per-value failure-mode battery:

  • hallucinated / off_target / incomplete / format Nouls, each framed so true = something is wrong;
  • the cookbook's dedicated omission question for values that came back empty (wrong only when the source supports a value the extractor missed);
  • per record, p_wrong is the max over its checks — never a mean, so one fired flag cannot be diluted by clean siblings — and any value at wrong_at (default 0.7, the cookbook's FIRE_T) escalates the whole audit.

Fail-closed throughout: a malformed answer marks the record invalid_response and escalates; a truncated source demotes pass to review. The state carries { purpose, source, records } with backticked path references in the questions, and every question gets the anti-injection framing.

The README's Multimodal intake section documents the full cascade this enables — Jev is text-only by contract, so images/scans/audio go: extract (host vision/ASR model: dense transcript + values) → jev_screen the transcript → jev_audit the values against it → judge with the existing tools. The cross-check is sound because the transcript and the values are independent passes over the same source.

Bounds: 32 records per call, 500 chars per request line, 2,000 per value, source at the shared 50,000 cap. Skill routing added (SKILL.md table + policy bullet, reference/tools.md block); tool-count assertions updated everywhere.

Testing

npm run typecheck, npm run build, npm test: 244 tests. New coverage: the one-request battery and scoped wire shape, clean pass, fabrication escalation without diluting clean siblings, the empty-value omission path, fail-closed on malformed answers, truncated-source demotion, and wrong_at as a parameter.

Notes

  • Question design is the cookbook's battery nearly verbatim, credited in the code comment and the README section, per CONTRIBUTING's cookbook-linking note.
  • No new dependencies; no vision model inside the server — the caller supplies both text artifacts, the tool owns the questions and the gate, the same split as jev_extract.

Jev reads text only, so multimodal intake lands as a cascade: the host
vision/ASR model produces a dense transcript and extracted values,
jev_screen checks the transcript, and the missing piece was a verifier
for the values themselves (jkudish#45).

jev_audit runs the SDE cascade cookbook's per-value failure-mode
battery (hallucinated / off_target / incomplete / format, each framed
true = something is wrong) in one request, with the cookbook's dedicated
omission question for empty values. p_wrong is the max over a value's
checks — never a mean, so one fired flag cannot be diluted by clean
siblings — and any value at wrong_at (default 0.7) escalates the whole
audit. Empty values get only the omission check; malformed answers mark
the record invalid_response and escalate; a truncated source demotes
pass to review. The README documents the full multimodal intake flow,
and the agent skill routes it.
@jkudish
jkudish merged commit 0ecb02f into jkudish:main Sep 28, 2026

jkudish commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Thank you for this — the proposal and implementation are both excellent, and the failure-mode battery is a strong fit as the twelfth tool. Merged with a few amendments during review:

  • The multimodal docs now describe the audit as a cross-check between two text artifacts, not a guarantee: one host model usually produces both the transcript and the values, so a shared misreading can pass. Prefer the original text when it exists.
  • The jev_screen step now says to honor block / review before passing the transcript onward, and that the screen protects your context, not the upstream model.
  • Added two live tests: planted fabrication / off-target / omission / format cases, and an embedded-injection transcript that must not buy a pass.

Shipping in the next release.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants