Skip to content

Repository files navigation

ScribeBench

License: MIT Data: CC-BY-4.0 Ranked: PriMock57 n=57 Tests: 114

A public source-vs-note QA workbench for AI-scribe claims.

Use the public site first: paste a source encounter and an AI-written note to check whether the note invented care, copy a reviewer-ready QA finding, or turn a broad vendor/model claim into the evidence it would actually need. The repo is the reproducible machinery behind that public surface: browser checker, optional model-backed Lab APIs, TypeScript evaluator, public cases, and scores-only evidence ledger.

The shortest version: people holding evidence use ScribeBench to catch care the AI note claims but the source does not support. Clinical QA reviewers, buyers, and builders leave with a QA finding first; bigger claims have to earn aggregate rows.

ScribeBench measures whether an AI-generated clinical note is faithful to the source encounter — it rewards capturing what the clinician said and did, and penalizes fabrication: invented findings, escalated diagnoses, workups that never happened.

The system loop is simple: one note becomes a QA finding; one claim becomes an evidence ask; many declared notes can become a scores-only public row. The repo contains the walk-up website, browser checker, model-backed Lab APIs, TypeScript evaluator, public cases, worklog, and evidence ledger needed to make that loop reproducible.

Use this repo in public

Start with the live artifact, not the file tree:

  1. Have a source encounter and an AI note? Open the one-note checker, paste both sides, and copy the QA finding.
  2. Have a vendor, model, or leaderboard claim? Open the claim checker and turn the claim into a public evidence ask.
  3. Want a current ranking? The honest answer is "not yet." The evidence ledger shows historical launch rows, smoke rows, and the current row blocker so nobody cites old rows as today's winner.

What the repo adds is reproducibility: the checker, cases, scoring code, worklog, and aggregate-only ledger behind those public artifacts.

What you can do today

  • Check one AI-scribe note in the browser and copy a reviewer-ready QA finding with excerpts, issue flags, boundaries, and the next proof step. The no-key checker is built for clinical QA reviewers, buyers, and builders with a source encounter plus generated note in hand.
  • Choose the smallest honest next step after a finding — review one note, challenge a larger claim, finish a current row, or explain the repo.
  • Challenge a claim such as "hallucination-free," "safe note," or "best model" and get the evidence ask that would make it public and testable.
  • Explain the repo as a public QA harness: browser checker, optional Lab APIs, evaluator, public cases, worklog, and scores-only evidence ledger.
  • Read the evidence ledger without mistaking historical launch baselines or smoke checks for a current model ranking.
  • Copy the citation boundary so old launch rows are discussed as failure-gradient evidence, not as today's best-model claim.
  • Help finish the current PriMock57 row by resuming the public API run or submitting a scores-only aggregate row from your own system.

Who this is for

  • Clinical QA reviewers paste one source encounter plus one AI-scribe note, then copy a hold/edit QA finding for review.
  • Buyers and operators paste a vendor or pilot claim, then copy the evidence ask that belongs in diligence.
  • AI-scribe builders reproduce unsupported-care failures, then route the defect with note/source excerpts.
  • Research contributors add current, scores-only aggregate rows instead of treating old baselines as a buying guide.

It is not a patient app, billing tool, clinical-clearance engine, or current-model winner board.

Start here

If you want to... Go here What you get
Check one AI-scribe note One-note checker A no-key, copy-ready QA finding with source-note issues, excerpts, evidence boundaries, and the next proof step.
Decide what to do after a finding After-the-finding guide Four bounded paths: hold or review the note, challenge a larger claim, finish a scored row, or explain the repo.
Challenge a vendor or model claim Claim checker A plain-language evidence ask for claims like "hallucination-free," "safe note," or "best model."
Bound what public claims can say Evidence ledger A QA-finding-first ledger that separates useful one-note output, historical rows, smoke tests, and current-model gaps.
Review the current blocker Current run status The live 9/57 scored blocker status, excluded-case boundary, and resume command for the public API row.
Add a citable public row Contribute aggregate evidence The scores-only evidence package, private candidate-note JSON shape, and benchmark command.
Reuse or audit the machinery site/, eval/, leaderboard/results.json The static site, evaluator, public cases, and scores-only ledger behind the public artifacts.

The public website is not a consumer app, a patient app, clinical clearance, or a current model buying guide. The existing rows are launch baselines and smoke checks; current claims need new powered rows.

Under the hood, ScribeBench adds two things to prior clinical-note evaluation work such as ACI-Bench, MEDIQA-Chat, and MedHallu:

  1. A fabrication tier that distinguishes registering delivered care from inventing what didn't happen. An ambient scribe's job is to register care the clinician actually delivered — a critical-care-time attestation, a placement statement captured from the encounter. That is the product working, not a hallucination. Other taxonomies flag it as one. ScribeBench tiers it as STANDARD and reserves DANGEROUS for content asserting something that did not happen. See docs/fabrication-taxonomy.md.

  2. An honest accounting of rater fragility. In the calibration work behind ScribeBench, three board-certified physicians reviewed the same 36 blind A/B note pairs (84 total ratings). The primary overlapping rater pair agreed at κ = 0.028 across 35 shared ratings (wide confidence interval), barely above chance. And in the same production data, binary structural completeness correlated with physician preference at ρ = −0.077 (not significant): the most complete note is not the one physicians prefer. If your eval rests on a single rater or a checklist, you are measuring the rater or the checklist, not quality. ScribeBench reports aggregate rates with bootstrap confidence intervals for this reason.

Companion preprint: Closed-Loop Quality Assurance for Production Clinical AI Documentation (medRxiv, DOI: 10.64898/2026.05.27.26353977v1). ScribeBench is its open walk-up artifact: one-note QA findings, aggregate evidence rows, and a public contribution path for testing whether AI-scribe notes stay faithful to the source. Disclosure: authored by a Sayvant co-founder; the judges and rubric are generalized from Sayvant's production QA system. See Disclosure.


Evidence ledger, not a current leaderboard

The historical table orders powered PriMock57 runs by dangerous-fabrication rate (lower is better), then narrative mean (higher is better), so the failure gradient is visible. It is not a current buying guide. Submit a current system via PR — see leaderboard/SUBMISSION.md.

Data policy: the leaderboard stores aggregate scores only — never raw model-generated note text. Full candidate notes are published only for open-weight models or your own runs. This respects provider output terms (publishing closed-model outputs as a redistributable dataset is not something we do). The bundled dataset is CC-BY synthetic + PriMock57 only.

The current ranked rows are historical launch baselines from June 2, 2026. They prove the powered PriMock57 path and show the failure gradient, but they are not a current buying guide. The next public work is to add current production, frontier, open-weight, and vendor-system rows as powered PriMock57 runs.

As of July 1, 2026, the production Vercel path has a public current-run blocker status, not a ranked current row: 30 PriMock57 cases were selected, 30 were attempted, 13 generated candidate notes, 9 scored, and 21 blocked or errored after the OpenRouter free-model cap was hit. The partial aggregate is useful plumbing evidence only; a citeable current row still needs at least 30 scored PriMock57 cases, preferably all 57, with declared repeats and exclusions.

System Dataset n Narrative ↑ (95% CI) Fidelity ↑ Dangerous-fab ↓ (95% CI) Leak ↓ Judge
claude-sonnet (scribe) PriMock57 57 78.4 [76.5, 80.4] 4.46 5.3% [0–12%] 0.0% claude-opus
gpt-4.1 (scribe) PriMock57 57 73.6 [72.3, 74.9] 4.27 5.3% [0–12%] 0.0% claude-opus
gpt-4o (scribe) PriMock57 57 67.4 [65.9, 68.7] 3.91 8.8% [2–18%] 0.0% claude-opus
claude-haiku (scribe) PriMock57 57 67.5 [65.7, 69.3] 4.22 24.6% [14–35%] 0.0% claude-opus

All are scores-only launch baselines (closed-model note text not published per the data policy) from a generic scribe prompt — not tuned production systems and not a current market ranking — judged by Claude Opus at repeats=2 over the 57 audio-grounded PriMock57 consults. The point is the failure gradient and the public harness: frontier launch models (claude-sonnet, gpt-4.1) fabricated dangerous content on ~5% of real consultations, gpt-4o on ~9%, and a small model (claude-haiku) on ~25%. Current model rows should be added as new powered PriMock57 runs.

Synthetic demo set (n=3, illustrative — wide CIs by design)
System n Narrative ↑ Fidelity ↑ Dangerous-fab ↓ Judge
openrouter-nemotron-3-ultra-live-smoke 3 100.0 5.00 0.0% nvidia/nemotron-3-super-120b-a12b-20230311:free
gpt-4o (scribe) 3 71.7 4.67 0.0% claude-opus
claude-sonnet (scribe) 3 68.0 4.67 33.3% claude-opus
example-baseline (seeded fab) † 3 59.0 3.50 33.3% claude-opus

The OpenRouter row was generated through the production Vercel API on June 30, 2026 with the current free-model path (n=3, repeats=1). It is useful as fresh plumbing evidence and deliberately not ranked. The production judge JSON-repair path was exercised on SYN-003, which is exactly the kind of operational fragility the public Lab should expose before anyone claims a system-level result.

SYN-003 carries a deliberate seeded fabrication. The claude-sonnet 33% is a real catch — on SYN-003 it fabricated "arrival via EMS" when the source says the daughter brought the patient in. At n=3 these CIs are wide on purpose; the PriMock57 table above is the substantive board.

Submit a current system to replace the launch baselines with evidence people can actually cite — see leaderboard/SUBMISSION.md.

Live results: leaderboard/results.json. Rows marked claimLevel: "smoke" are visible for transparency but are not ranked.

Public website

This repo builds a static public ScribeBench site for Vercel, currently live at https://scribe-bench.vercel.app.

Use the website for these jobs:

  1. Choose the public job. The first screen now says the point, user, and reason for the repo before routing three visitor situations: "I have source + note," "I heard a claim," and "I can add proof." On mobile, the source-note route leads to the paste form before the seeded demo, keeping ScribeBench grounded in the visitor's note review instead of a vague leaderboard tour.
  2. Check one note. Paste a source encounter and an AI-written note into the source-note intake. The browser-only checker returns a reviewer handoff, a copied QA finding that starts with Decision, Action, Why, Evidence, and Next, and a routing note for chart QA, builder/vendor defects, and public claim boundaries. It catches high-confidence invented care and chart-fact drift: unsupported treatments, procedures, medication changes, diagnoses, orders, test results, demographics, laterality, allergies, and leaks. No API key required.
  3. Challenge a claim. Turn "hallucination-free," "safe note," "best model," or similar language into the evidence level it would actually require.
  4. Read today's answer. The checker, after-the-finding guide, and evidence ledger now put the one-note QA finding before the aggregate-row work: what ScribeBench can support today, what the current PriMock57 row is still missing, and why old rows should not be cited as a current winner board.
  5. Choose the next bounded action. After a finding, decide whether to review the note, challenge a broader claim, finish the current row, or explain the repo.
  6. Run optional Lab checks. Use configured provider models for generation and second-opinion judging when you want plumbing evidence. Treat those second-read reviews as smoke or review artifacts, not ranked results.
  7. Publish aggregate evidence. Use the run builder and submission path to add aggregate PriMock57 or real-workflow rows without publishing raw closed-model notes.

The site is not the scribe product, not a patient app, not clinical clearance, and not a current model buying guide. The historical launch rows prove the harness and failure gradient; current claims need new powered rows.

What the repo contains:

Piece Files Purpose
Public website site/ Static Vercel site with the one-note checker, claim checker, evidence ledger, optional second-opinion Lab, and aggregate row builder.
Browser checker site/local_receipt.js No-key source-vs-note triage that turns one note into a copy-ready QA finding with unsupported-care flags, excerpts, evidence boundary, and next ask.
Live API api/generate.js, api/judge.js, api/models.js Optional model-backed generation and judging for the Lab through OpenRouter or Baseten-compatible APIs.
Eval engine eval/ TypeScript harness for narrative quality, input fidelity, dangerous fabrication, leak checks, repeats, and bootstrap intervals.
Data data/synthetic/, data/primock57/ Synthetic demos plus 57 public PriMock57 consults. No real patient data.
Evidence ledger leaderboard/results.json, site/current-run.json, site/worklog.json Scores-only public rows, current-run status, and build-in-public changelog.
npm run build
npm run preview

Source lives in site/. The build script copies the static app into dist/ and publishes bounded JSON from the benchmark artifacts.


Metrics

Metric Direction What it measures
Narrative mean higher 6-dimension physician-style quality (0–100), calibrated to blind preference
Input fidelity higher Does the note faithfully capture what the clinician said? (1–5)
Dangerous-fabrication rate lower Fraction of notes that assert something that did not happen — invented findings, escalated diagnoses, contradicting orders
Leak rate lower Fraction of notes containing raw template placeholders or internal-metadata tokens (deterministic, 0-token scan)

The fabrication judge draws a line most metrics miss: registering care the clinician actually delivered is not fabrication — a critical-care-time attestation or a placement statement captured from the encounter is the value of an ambient scribe, not a hallucination. Only content asserting something that did not happen is dangerous. See docs/methodology.md.


Quickstart

npm install

# 1. Smoke the harness with bundled synthetic data. This is not a ranked claim.
npx tsx eval/run_benchmark.ts \
  --dataset data/synthetic/cases \
  --candidate data/synthetic/example_candidate.json \
  --system "example-baseline" \
  --out leaderboard/_smoke-pending.json

# 2. For a real row, run your own scribe/system over PriMock57 first.
# Save candidate notes as JSON:
#   [{ "caseId": "PM57-d1c01", "note": "..." }, ...]
# Keep raw closed-model notes local unless provider terms allow publication.

# 3. Score that candidate file with a separate judge backend:
#   anthropic  — needs ANTHROPIC_API_KEY
#   cli        — uses the `claude` CLI over OAuth (Max/Pro plan, no key)
#   baseten    — OpenAI-compatible Baseten Model APIs, needs BASETEN_API_KEY
#   openrouter — OpenAI-compatible OpenRouter, needs OPENROUTER_API_KEY
export SCRIBEBENCH_BACKEND=baseten
export BASETEN_API_KEY=...
export SCRIBEBENCH_JUDGE_MODEL=declared-strong-judge-model

npx tsx eval/run_benchmark.ts \
  --dataset data/primock57/cases \
  --candidate /tmp/current_system_primock57_notes.json \
  --system "current-system-under-test" \
  --repeats 2 \
  --out leaderboard/_pending.json

The example candidate deliberately seeds one fabrication (case SYN-003 invents a head CT and a syncope workup the source rules out) so you can see the fabrication judge fire. It is useful for plumbing and demos, not for ranking systems.

If you want ScribeBench to generate candidate notes for a smoke run, use scripts/generate_baseline.ts with --gen openrouter or --gen baseten; treat that as a helper, not the default meaning of the benchmark. The publishable path is a declared candidate-note file from the actual system you want to discuss.

Current public-API run path

When the live Vercel site has provider keys configured, you can add a current PriMock57 evidence row through the same public APIs visitors use. This keeps raw generated notes in the ignored local cache and writes an aggregate pending file:

npm run bench:public-api -- \
  --base-url https://scribe-bench.vercel.app \
  --dataset data/primock57/cases \
  --system current-system-public-api \
  --repeats 1 \
  --out leaderboard/_public-api-pending.json \
  --status-out site/current-run.json

By default, the runner forwards a local OPENROUTER_API_KEY or BASETEN_API_KEY to the public API as the matching temporary provider header. Use --key-env MY_OPENROUTER_KEY to point at a different environment variable. The key is never written to the progress cache or pending artifact.

The progress cache lives under .scribebench-cache/public-api-runs/ so interrupted runs can resume. The runner also writes site/current-run.json for PriMock57 attempts by default, so every retry updates the public status card with the latest selected/generated/scored/errored counts and blocker. Do not copy a pending row into leaderboard/results.json until the run has enough completed PriMock57 cases, no unreviewed errors, declared model/judge details, and the self-judge/second-judge limitation is disclosed. Free OpenRouter models can be slow and quota-limited; when a judge call times out or hits the daily cap, the runner records the case as errored/excluded and can resume once credits or another judge backend are available. The public site exposes that blocker in /assets/current-run.json and gives a copyable resume command in the Evidence section.

Current status: on July 1, 2026 ICT (July 1, 2026 03:04 UTC), the live public API runner selected and attempted 30 PriMock57 cases. It generated notes for 13/30, scored 9/30 attempted cases (9/57 target cases), and left 21/30 blocked or errored after the OpenRouter free-model cap was hit. This is a blocker status and partial plumbing signal, not a model result or current ranking. Resume with credits, a non-capped provider key, or a second judge before publishing any ranked current row.

Datasets

  • data/synthetic/cases/ — 3 fully synthetic encounters (ED, clinic, inpatient). Ships in-repo; the runnable quickstart set.
  • data/specialty/cases/ — synthetic emergency-medicine + hospital-admission cases (physician-authored), broadening coverage beyond primary care. Extensible — PRs welcome.
  • data/primock57/cases/ — 57 audio-grounded mock primary-care consultations derived from PriMock57 (Babylon Health, CC-BY-4.0), vendored as ScribeBench cases. Regenerate from upstream with bash scripts/fetch_primock57.sh && npx tsx scripts/build_primock57_cases.ts. This is the scored leaderboard set.

ScribeBench contains no real patient data. Synthetic or already-public, appropriately licensed corpora only. See CONTRIBUTING.md.

Bring your own pipeline

The eval engine is a small, dependency-light TypeScript library:

  • eval/narrative_judge.tsevaluateNarrative(note, { source })
  • eval/fabrication.tsjudgeFabrication(note, source) + detectLeaks(surfaces) (pure, no LLM)
  • eval/llm.ts — pluggable backend (Anthropic API / Claude CLI / Baseten / OpenRouter / your own)

See docs/model-backends.md for Baseten, OpenRouter, and powered-run examples.

Prior work

ScribeBench builds on a line of clinical-note evaluation work and is explicit about what it adds:

  • ACI-Bench (Yim et al.) — ambient clinical-note generation with a hallucination + omission taxonomy (major/minor severity). The closest prior art. ScribeBench differs in the delivered-care-vs-invented fabrication tier and in physician-preference calibration.
  • MEDIQA-Chat / MEDIQA-Sum — clinical dialogue-to-note shared tasks.
  • MedHallu, MedHallBench — medical-LLM hallucination benchmarks (mostly QA, not note generation).
  • npj Digital Medicine (2025) — a clinical-safety/hallucination framework for LLM medical summarization.

What ScribeBench adds: (1) the fabrication tier that does not penalize registering delivered care; (2) physician-preference calibration with an honest rater-fragility result (κ); (3) a fidelity-first, scores-only leaderboard.

Disclosure

ScribeBench is authored by a Sayvant co-founder. Sayvant builds clinical documentation AI and competes with vendors who may appear on this leaderboard. The judges and rubric are generalized from Sayvant's production QA system (no production prompts, no patient data — see docs/methodology.md). To keep the benchmark neutral:

  • Sayvant does not submit its own leaderboard row at launch. If it ever does, it will be clearly flagged and scored under an independent judge configuration.
  • The eval engine, rubric, and synthetic data are fully open so anyone can audit or re-run the scoring.
  • Prior work is credited above; the novel contributions are scoped narrowly and honestly.

License

  • Code: MIT (LICENSE)
  • Synthetic data + rubrics: CC-BY-4.0 (LICENSE-DATA)
  • PriMock57 retains its upstream CC-BY-4.0 license.

Citation

See CITATION.cff.

About

Public workbench for checking whether AI-scribe notes invent care, with receipts and aggregate evidence rows.

Topics

Resources

Contributing

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages