An evaluation harness that audits AI classifiers and agents — catching the mistakes a plain accuracy score hides.
Most AI projects end with one number: "my model is 94% accurate." That number sounds like an answer, but it rarely tells you the thing you actually need to know — accurate at what, and what happens on the 6% it gets wrong?
Errata is not another classifier. It's the tool that sits on top of any classifier or agent and checks whether its score can actually be trusted. Hand it a model and a test set, and instead of one number, it hands back an itemised record: which specific mistakes the model made, how far off each one was, what it would actually cost, whether the model's confidence can be trusted, and whether it can be tricked.
Think of it less like a student taking an exam, and more like the person standing at the door afterward, checking whether the exam actually proved anything.
An errata is the corrected record a publisher attaches to a book after it's already printed — not a star rating, but an itemised list of exactly what's wrong and where. This project does the same thing for a model's predictions: instead of one clean verdict, it hands back the itemised correction record. See docs/research-notes.md for the full reasoning, including the honest caveat on where that metaphor doesn't perfectly apply.
The short version:
- Start with a test set where the right answers are already known — but keep them hidden from the model while it makes its guesses.
- Let the model predict, blind. Whatever it hands back (a label, a score, a number, or free text) gets standardised by a small adapter layer before anything else touches it.
- Only after every guess is locked in, reveal the real answers and compare.
- Score the result several different ways at once: a plain flat accuracy score, how far off the guess was on the category tree, what it would cost in rupees, and whether the model's stated confidence matched how often it was actually right.
- Mix in a handful of deliberately adversarial test cases, and separately re-run the whole thing while dialling the rare-case rate up or down.
- Put it all together into one report — not one number pretending to speak for everything.
Full step-by-step pseudocode for the harness itself is in docs/pseudocode.md.
This repo actually holds two different projects at two different stages, and it's easy to mix them up if you don't know that going in:
| Errata itself (the harness) | The base model (the agent being evaluated) | |
|---|---|---|
| What it is | The evaluation tool — plain logic, no ML, the "auditor" | A vendor-payment fraud triage agent — the first real subject Errata will eventually evaluate |
| Status | Research, design, pseudocode, and a proven dry run — done. Real Python implementation — not started. | Bayesian reasoning, decision policy, and a working Stage 8 simulation — done and executed. |
| Where it lives | docs/, src/errata/, dry_run_demo.py |
base-model/ |
| Its own flow diagram | assets/errata-flow-diagram.png (shown above) |
base-model/assets/base-model-flow-diagram.png (shown below) |
| Read this first | docs/research-notes.md | base-model/decision/probability-decision-record.md |
Why the base model exists at all: Errata needs something real to evaluate once it's built — a model or agent with actual decisions, actual probabilities, actual mistakes to catch. Rather than wait until Errata's code is finished to pick that subject, the fraud triage agent was built and reasoned through first, so there's a concrete, already-tested agent ready and waiting the moment Errata's own scoring logic is implemented.
This is not Errata — this is the fraud-triage agent Errata will eventually evaluate. Follow the numbered steps ①→②→③→④ across the top: one "please change our bank account" email comes in, evidence is collected, a Bayesian belief update runs across five hidden causes, and a three-action decision policy (Pay / Verify / Escalate) picks exactly one action. The section below the dashed line is how that agent is tested offline — 1,000 synthetic cases with a known answer, graded, scored, and stress-tested under five what-if scenarios. Full reasoning behind every number in this diagram is in base-model/decision/probability-decision-record.md.
Two pieces of this don't exist anywhere as a ready-to-use tool, as far as the research in this repo could find:
- Tree-distance scoring — treating "mistook phishing for malware" as a smaller error than "mistook phishing for safe." The underlying idea is published research (Apple's Neo, CHI 2022), but no open, installable implementation of it exists — this repo builds it from scratch.
- Cost + hierarchy + calibration combined in one report — most tools do one of these in isolation. Combining "how severe," "how expensive," and "was the confidence honest" into a single evaluation is the actual contribution here.
Full literature review and gap analysis: docs/research-notes.md.
| Piece | Status |
|---|---|
| Problem research & prior-art check (Errata) | Done — see docs/research-notes.md |
| Tools & architecture decisions (Errata) | Done — see docs/tools-and-approach.md |
| Full pseudocode, 9 steps (Errata) | Done — see docs/pseudocode.md |
| Flow diagram (Errata) | Done — see assets/errata-flow-diagram.png |
| Dry run proving the scoring logic works (Errata) | Done — see docs/dry-run-walkthrough.md |
| Actual Python implementation (Errata) | Not started |
| Base model — priors, evidence, Bayesian update | Done — see base-model/decision/probability-decision-record.md |
| Base model — cost-based decision policy & thresholds | Done — see base-model/decision/probability-decision-record.md §7 |
| Base model — pseudocode | Done — see base-model/pseudocode.md |
| Base model — flow diagram | Done — see base-model/assets/base-model-flow-diagram.png |
| Base model — Stage 8 simulation (executed, not just designed) | Done — see base-model/experiments/stage8-fraud-triage-simulation/ |
| IJCAI-format preprint (paper covering both Errata and the base model) | Done — see paper/preprint.pdf |
| Community validation / discussion | Ongoing — see links below |
Posted for genuine outside feedback on the problem framing, the Bayesian reasoning, and whether this gap in evaluation tooling is real:
- r/AskStatistics — sanity-checking the entropy-increased finding
- r/cybersecurity — reality-checking the fraud / verification assumptions
- r/learnmachinelearning — the project write-up and what surprised me building it
- r/AI_Agents — how people actually evaluate agent decisions, not just outputs
Each thread and what came out of it is logged in discussion-record.md.
errata/
├── README.md ← you are here
├── LICENSE
├── requirements.txt
├── .gitignore
├── .env.example
├── dry_run_demo.py ← proves Errata's scoring logic works, no real model needed
│
├── base-model/ ← the AGENT being evaluated (not Errata itself)
│ ├── pseudocode.md ← this agent's own decision-logic pseudocode
│ ├── assets/
│ │ └── base-model-flow-diagram.png ← THIS agent's own diagram, separate from Errata's
│ ├── decision/
│ │ └── probability-decision-record.md ← priors, evidence, Bayes update, cost thresholds
│ └── experiments/
│ └── stage8-fraud-triage-simulation/
│ ├── fraud_triage_simulation.py ← runs the simulation + 5 stress-test scenarios
│ └── results/
│ ├── summary.md ← analysed results, findings, verdict
│ └── raw-output.txt ← unprocessed console output
│
├── docs/ ← ERRATA'S OWN docs (the harness, not the base model)
│ ├── research-notes.md ← the "why" behind Errata, prior-art check
│ ├── tools-and-approach.md ← tools, decisions, and reasoning
│ ├── pseudocode.md ← Errata's own step-by-step logic, all 9 steps
│ └── dry-run-walkthrough.md ← one case traced by hand + full batch output
│
├── assets/
│ └── errata-flow-diagram.png ← ERRATA'S OWN diagram, separate from the base model's
│
├── paper/
│ ├── main.tex ← IJCAI-format preprint source
│ └── preprint.pdf ← compiled paper covering both Errata and the base model
│
├── src/
│ └── errata/ ← Errata's actual implementation goes here (not started yet)
│
└── tests/ ← tests for Errata's own scoring logic
Quick rule of thumb: if you're looking for the harness — how Errata scores things, what makes it novel, its own design and its own diagram — go to docs/ or assets/errata-flow-diagram.png. If you're looking for the agent being evaluated — its priors, its Bayesian reasoning, its decision policy, its own diagram, or the simulation proving that policy works — go to base-model/.
No real model or setup needed — this proves Errata's scoring logic itself works, using hand-mocked predictions:
python dry_run_demo.py
See docs/dry-run-walkthrough.md for the full trace and what the output actually proves.
No setup beyond the standard library — this runs the fraud-triage agent's decision policy across 1,000 synthetic cases plus five stress-test scenarios:
cd base-model/experiments/stage8-fraud-triage-simulation
python fraud_triage_simulation.py
See base-model/experiments/stage8-fraud-triage-simulation/results/summary.md for the analysed findings, or the decision record's §11 for the full reasoning behind the test design.
The full IJCAI-format preprint covering both the base model's Bayesian design / experiment and Errata's own evaluation methodology (including where the two currently do — and don't yet — connect) is in paper/preprint.pdf.
git clone https://github.com/<your-username>/errata.git
cd errata
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
pip install -r requirements.txt
Copy .env.example to .env and fill in any local settings (e.g. your LiteLLM/Ollama endpoint) before running anything that needs a model.
MIT — see LICENSE.
Built by Arif Hussain as a personal project to demonstrate evaluation rigor for AI systems — not just building a model, but proving whether one is actually safe to trust.

