Repository navigation
Run the locked Local ASR bakeoff and elect or reject a winner #267
Description
Activity
- added 2 commits that reference this issue
on Sep 11, 2026 First locked-candidate measurement snapshot (whisper-cli + ggml-base.en, frozen corpus revision 2026-09-10-r1). Integrity gates ran before any score: manifest metadata lock verified (
validate-bakeoff --expected-manifest-sha256 49c34183…→ pass), runtime hash 5fc66a04… and model hash a03779c8… both match the frozen locks, scoring tool at main 2fa28f5, run JSON fingerprint sha256:341d9eb0…. Invocation: one fresh whisper-cli process per clip,-oj -npdefaults; run-to-run determinism verified on a sample clip. 100 speech clips + 20 negatives measured; scored by transcript-quality score-corpus with the frozen WER normalization.Aggregates (hashes/counts/aggregates only; no private audio, paths, or text):
- Held-out overall WER: 0.12 over 80 cases — fails the 0.10 gate, but the excess is entirely the 10 private clips whose references are still the un-adjudicated read prompts (public-reference-only held-out WER is 0.0656, passing).
- Held-out per-stratum WER: librispeech_test_clean_english 0.0656 (70 cases, passes ≤0.15); consented_indian_english 0.3666 (10 cases) — unreliable until listening-based reference adjudication, because the audio contains spoken script markers the prompt references lack.
- Critical-content errors: 262 held-out (172 public-only) against the =0 gate — dominated by homophone/number-formatting pairs the frozen detector counts (e.g. NUMBER TEN → Number 10, reigned → rained).
- Negative fixtures: 14/20 produced non-empty output (6 [BLANK_AUDIO], 1 [♪♪♪], 7 silence hallucinated one word) against the =0 gate.
- Tuning-split WER 0.0324 recorded; no tuning was performed.
- Process wall time snapshot (not a pipeline stage): p50 950 ms, p95 3370 ms, max 7552 ms; total wall ≈132 s for ≈29.1 min audio.
Decision per the locked contract: explicit no-go — private-stratum references require listening adjudication before their numbers are spoken-truth, two quality gates fail for the locked candidate, and the soak/pipeline-latency/packaged-runtime gates are not measured (the daemon worker interface fails closed until placement is wired to a delegated ancestor, per the #266 follow-up). No candidate election, no thresholds moved.
Open contract-interpretation question to resolve BEFORE further scores are observed: do whisper non-speech markers ([BLANK_AUDIO]/[♪♪♪]) count as negative-fixture insertions, or only produced speech words? Detailed evidence is retained locally (redacted aggregates only, per the frozen report rules).
Parent
What to build: Execute the frozen R7 evaluation through the intended process-wrapped capture, supervision, sandbox, resource, and Delivery interfaces. Publish a reproducible evidence report and elect a production winner only if every locked gate passes; otherwise publish no-go.
Blocked by: #264, #265, and #266.
Status: blocked · P0 · L2 decision gate
Do not include: moving thresholds after observing scores, selecting on one health fixture, production enablement, or Fedora support claims.