feat: add evidence-refined AV ask and truthful usage receipts - #2
Conversation
|
Runtime review closure at tested head 75125c3:
Verification: 267 offline tests passed; source distribution and wheel built successfully; git diff --check passed. No live model calls were made. |
|
Independent runtime re-review of commit The two previously identified P2 defects are resolved:
Also verified that answer generation sends the configured Validation at this exact commit: No open P1/P2 findings in this runtime review scope. Rejection handling, broad-summary isolation, successful-frame citation validation, opt-in fallbacks, and incomplete-usage behavior were included in the review. This review makes no claim about live model accuracy or measured cost savings; those require the separate reproduction receipts. |
|
[P2] Natural-language questions can return no candidates before refinement runs Confirmed at
Please add a generic natural-language candidate query path for This finding extends the earlier review beyond the four fixes; it leaves one P2 open at this head. |
|
Implemented the natural-language retrieval fix in
Offline regressions cover the generic truck question with and without punctuation, Unicode/punctuation, video isolation, stopword-only and unrelated questions, term limits, explicit FTS queries, original-question embeddings, and Jev rejection of broad candidates. Validation before commit: 278 tests passed across the full offline suite; |
|
Independent review of natural-language candidate retrieval at exact head Verified that Independent validation: |
|
PASS — independent review of exact head No actionable findings. Offline verification:
This is a review comment only; I did not approve or merge. |
|
PASS — independent re-review of exact head 7e1a0ce, focused on 86b849d..7e1a0ce. All four prior findings are resolved:
Validation at the exact detached head: 16 cost-model tests, 5 transcript-helper tests, 31 provider/config tests, notebook check, full 283-test suite, and sdist/wheel build all passed. GitHub CI is green on Python 3.11–3.13. No API calls, approval, merge, or publication occurred. This is a review comment only; the PR remains draft. |
Independent exact-head review — PASSReviewed PR #2 at
Provenance limitations: costs are cache-aware list-rate estimates, not provider invoices; infrastructure/storage/network costs remain unknown. Review did not re-run model or provider traffic. |
|
STATE — fresh authorized live attempt recorded at |
|
STATE — live Jev comparison arm closed by credential absence; no Jev request was made. Done:
Blocker:
Next:
|
Independent exact-head review — PASSReviewed PR #2 at
Provenance: dollar figures are cache-aware public list-rate estimates, not invoices; provider/proxy billing and infrastructure/storage/network costs remain unknown. Full tests/build were reported as passing by the operator and were not repeated in this independent checkout. |
|
Exact-head review — PASS Head: dad9dcf Offline validation in a detached read-only checkout:
New receipt checks:
Provenance limitation: this is an offline exact-head audit of recorded receipts and docs; external provider responses were not replayed or billed. |
|
Independent evidence review of [P1] Correct the Jev input-rate unit before deriving query and cumulative totals. In the refined-query receipt, relevance has 1,888 input tokens and an estimate of The Grok answer's cache-aware list estimate of Propagate the unit correction to scenario.receipts.json, the receipt's cumulative total, the cost-model README, and generated notebook/output. Preserve the raw token counts and failed-attempt accounting. Add a regression that uses 1 billion Jev input tokens and expects $42, so arithmetic-only tests cannot miss the rate-unit mistake. No new paid inference is necessary to fix these derived estimates. |
|
Implemented the Jev pricing correction at The unchanged 1,888 relevance and 1,464 support input tokens produce $0.000140784 in Jev list-price estimates. Adding the unchanged $0.00165685 Grok answer estimate gives $0.001797634 for the query's model calls. The known experiment estimate is $0.902456134, with $4.097543866 of estimate headroom under the $5 cap. These are not verified billed dollars. The receipt retains the original derived values and a dated correction. Measured usage, answer content and timings are unchanged. The regression derives the price from the published billion-token unit, includes free output, and checks both stages against the query and cumulative ledger. The notebook was regenerated. Related provenance now explicitly states that the answer is a normalized factual excerpt, Gemini implicit cache usage is unknown, and the ingestion timer excludes earlier frame extraction and external ASR. Validation: 283 AV tests and 18 cost-accounting tests passed; notebook consistency and diff checks passed. No provider requests were made. Independent review of this exact commit has been requested. |
|
Independent review at exact head The correction now derives TypeSafe's published $42 per billion input tokens = $0.042 per million, with free output, from the official model documentation. The query receipt reconciles:
The original $0.14244085 query estimate is retained as correction history. The $0.140643216 reduction reconciles the experiment ledger and generated notebook to $0.902456134 in known list-price estimates, with $4.097543866 estimate headroom under the $5 cap. Raw measured usage and timings were preserved; the four correction-recorded raw receipt checksums still match. Actual billing, infrastructure allocation, and the unmetered failed request remain unknown. The completed one-time preparation is $0.5048525 captions + $0.0685709 ASR = $0.5734234. The first AV question including that preparation is $0.575221034, versus the native Gemini receipt's $0.308076 full-rate estimate. These are model-call estimates, not complete service costs or invoices. The receipt/notebook now correctly leave Gemini's unreported implicit cache usage unknown. The answer is correctly labeled a normalized factual excerpt, not the complete verbatim returned response. Its core 12.2% manufacturing-growth answer agrees with the recorded transcript evidence and native response; the AV citation is the coarse 19:00–20:00 transcript window. This does not validate every omitted timestamp/visual statement, establish representative accuracy, or turn the Jev support score into a correctness probability. The reviewed runtime remains unchanged by this correction. In rag.py, no hits and successful evidence rejection stop without a supported answer; a Jev outage is explicitly marked raw/unjudged. Weak or unknown support can use the explicitly configured inspector, and its result is support-judged again. Inspection has bounded sampled-still budgets; it does not send native video/audio or establish events between frames. No stronger inspection occurred in the live query, so its quality benefit remains unmeasured. Independent offline checks on this exact head: 18 cost-accounting tests plus 12 subtests passed; 37 focused AV refinement/ingestion tests passed; deterministic notebook check and |
|
Independent exact-head residual review — PASS Head:
PASS for the residual code-correctness and disclosure-boundary review at this exact head. CI 3.11/3.12/3.13 is green. |
av askpreviously answered directly from raw retrieval. This change adds an optional Jev evidence-refinement path: retrieve candidates, judge relevance, bound and merge temporal context, generate a cited answer, then judge whether the evidence supports that answer. Refinement activates when a TypeSafe key is configured, with request/config opt-outs. Low-support answers can optionally use bounded sampled-frame inspection; that fallback was not exercised by the included live run.The implementation keeps successful rejection separate from provider failure: rejected evidence produces no supported answer, while a refinement outage warns before using raw retrieval. It also fixes natural-language FTS candidate selection, adds validated transcript-sidecar import, configurable provider token-limit fields, and per-stage usage/failure metadata. A provider ignoring an output cap remains visible in the receipts.
The cookbook now lives in AV and includes a standalone Gemini ASR sidecar script, reproducible commands/settings, sanitized receipts, and an offline cost calculator/notebook. One 75-minute source was ingested into 300 Grok frame captions plus 75 transcript windows, then queried using
grok-4.20-0309-non-reasoningwith two Jev judgments. A separategemini-3.8-flashnative-video/audio request used the unchanged question.The 2026-09-19 pricing correction derives Jev's official $42/billion input rate as $0.042/million with free output. The corrected query model-call estimate is $0.001797634, including the Grok answer; known experiment estimates including successful, failed, and trial calls total $0.902456134. The receipt preserves the earlier derived values and dated correction while keeping measured usage and timings unchanged. These are published-list-price estimates, not verified bills.
Validation on the correction:
git diff --checkpassed.Limits: the live evidence covers one question and establishes no aggregate accuracy, speed, or cost parity. Grok query latency excludes ingestion; the ingestion timer excludes earlier frame extraction and external ASR. The two pipelines used different sampling/input paths. No explicit Gemini cache was created, while implicit cached usage remains unknown. Actual provider/proxy billing and compute, storage, and network allocations remain unknown; complete-cost ratios stay suppressed. The sanitized Jev answer is a labeled normalized factual excerpt. No source media, caption/transcript corpus, credentials, upload identifiers, or private routes are included.