Skip to content

Run the locked Local ASR bakeoff and elect or reject a winner #267

Description

@Anuraj-dev

Parent

What to build: Execute the frozen R7 evaluation through the intended process-wrapped capture, supervision, sandbox, resource, and Delivery interfaces. Publish a reproducible evidence report and elect a production winner only if every locked gate passes; otherwise publish no-go.

Blocked by: #264, #265, and #266.

Status: blocked · P0 · L2 decision gate

  • Revalidate primary model cards, licenses, immutable hashes, native dependencies, runtime formats, and redistribution terms.
  • Evaluate whisper.cpp first; evaluate another candidate only if its exact runtime/package stack can satisfy R3-R5.
  • Run all tuning, held-out, negative, duration-stratum, and critical-content cases from the frozen corpus without split leakage.
  • Report raw counts plus WER, punctuation/semantic errors, p50/p95/max stage timings, observed CPU/GPU, RSS/cache limits, and every rejected case.
  • Complete at least 200 consecutive Recordings over at least two hours, including 20 warmups, with no crash, duplicate Delivery, stale result, unexpected network attempt, unreaped child, or RSS growth beyond the locked bound.
  • Verify the exact packaged runtime under the intended service sandbox, not a developer-shell substitute.
  • Preserve private audio and detailed traces locally; publish redacted aggregates and immutable evidence identifiers.
  • Name the result Bakeoff Winner only if every frozen gate passes. Missing or failed evidence records an explicit no-go.
  • Obtain independent review of corpus integrity, assertion quality, measurements, and the election decision.

Do not include: moving thresholds after observing scores, selecting on one health fixture, production enablement, or Fedora support claims.

Activity

  1. Anuraj-dev commented on Sep 11, 2026

    @Anuraj-dev
    OwnerAuthor

    First locked-candidate measurement snapshot (whisper-cli + ggml-base.en, frozen corpus revision 2026-09-10-r1). Integrity gates ran before any score: manifest metadata lock verified (validate-bakeoff --expected-manifest-sha256 49c34183… → pass), runtime hash 5fc66a04… and model hash a03779c8… both match the frozen locks, scoring tool at main 2fa28f5, run JSON fingerprint sha256:341d9eb0…. Invocation: one fresh whisper-cli process per clip, -oj -np defaults; run-to-run determinism verified on a sample clip. 100 speech clips + 20 negatives measured; scored by transcript-quality score-corpus with the frozen WER normalization.

    Aggregates (hashes/counts/aggregates only; no private audio, paths, or text):

    • Held-out overall WER: 0.12 over 80 cases — fails the 0.10 gate, but the excess is entirely the 10 private clips whose references are still the un-adjudicated read prompts (public-reference-only held-out WER is 0.0656, passing).
    • Held-out per-stratum WER: librispeech_test_clean_english 0.0656 (70 cases, passes ≤0.15); consented_indian_english 0.3666 (10 cases) — unreliable until listening-based reference adjudication, because the audio contains spoken script markers the prompt references lack.
    • Critical-content errors: 262 held-out (172 public-only) against the =0 gate — dominated by homophone/number-formatting pairs the frozen detector counts (e.g. NUMBER TEN → Number 10, reigned → rained).
    • Negative fixtures: 14/20 produced non-empty output (6 [BLANK_AUDIO], 1 [♪♪♪], 7 silence hallucinated one word) against the =0 gate.
    • Tuning-split WER 0.0324 recorded; no tuning was performed.
    • Process wall time snapshot (not a pipeline stage): p50 950 ms, p95 3370 ms, max 7552 ms; total wall ≈132 s for ≈29.1 min audio.

    Decision per the locked contract: explicit no-go — private-stratum references require listening adjudication before their numbers are spoken-truth, two quality gates fail for the locked candidate, and the soak/pipeline-latency/packaged-runtime gates are not measured (the daemon worker interface fails closed until placement is wired to a delegated ancestor, per the #266 follow-up). No candidate election, no thresholds moved.

    Open contract-interpretation question to resolve BEFORE further scores are observed: do whisper non-speech markers ([BLANK_AUDIO]/[♪♪♪]) count as negative-fixture insertions, or only produced speech words? Detailed evidence is retained locally (redacted aggregates only, per the frozen report rules).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions