Skip to content

feat: add evidence-refined AV ask and truthful usage receipts - #2

Merged
seanphan merged 12 commits into
mainfrom
codex/jev-query-cascade
Sep 19, 2026
Merged

seanphan merged 12 commits into
mainfrom
codex/jev-query-cascade

Conversation

@seanphan

@seanphan seanphan commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

av ask previously answered directly from raw retrieval. This change adds an optional Jev evidence-refinement path: retrieve candidates, judge relevance, bound and merge temporal context, generate a cited answer, then judge whether the evidence supports that answer. Refinement activates when a TypeSafe key is configured, with request/config opt-outs. Low-support answers can optionally use bounded sampled-frame inspection; that fallback was not exercised by the included live run.

The implementation keeps successful rejection separate from provider failure: rejected evidence produces no supported answer, while a refinement outage warns before using raw retrieval. It also fixes natural-language FTS candidate selection, adds validated transcript-sidecar import, configurable provider token-limit fields, and per-stage usage/failure metadata. A provider ignoring an output cap remains visible in the receipts.

The cookbook now lives in AV and includes a standalone Gemini ASR sidecar script, reproducible commands/settings, sanitized receipts, and an offline cost calculator/notebook. One 75-minute source was ingested into 300 Grok frame captions plus 75 transcript windows, then queried using grok-4.20-0309-non-reasoning with two Jev judgments. A separate gemini-3.8-flash native-video/audio request used the unchanged question.

The 2026-09-19 pricing correction derives Jev's official $42/billion input rate as $0.042/million with free output. The corrected query model-call estimate is $0.001797634, including the Grok answer; known experiment estimates including successful, failed, and trial calls total $0.902456134. The receipt preserves the earlier derived values and dated correction while keeping measured usage and timings unchanged. These are published-list-price estimates, not verified bills.

Validation on the correction:

  • 283 existing AV tests passed.
  • 18 offline cost-accounting tests passed, including published billion-token unit conversion, free output, receipt/ledger reconciliation, missing cache usage, and reservation accounting.
  • Deterministic notebook regeneration/check and git diff --check passed.
  • No provider requests were made for the correction.

Limits: the live evidence covers one question and establishes no aggregate accuracy, speed, or cost parity. Grok query latency excludes ingestion; the ingestion timer excludes earlier frame extraction and external ASR. The two pipelines used different sampling/input paths. No explicit Gemini cache was created, while implicit cached usage remains unknown. Actual provider/proxy billing and compute, storage, and network allocations remain unknown; complete-cost ratios stay suppressed. The sanitized Jev answer is a labeled normalized factual excerpt. No source media, caption/transcript corpus, credentials, upload identifiers, or private routes are included.

@seanphan

Copy link
Copy Markdown
Contributor Author

Runtime review closure at tested head 75125c3:

  1. Answer generation now has a positive configurable output cap, passes max_tokens to the OpenAI-compatible request, appears in config show and ask_settings, and has an offline request regression.
  2. Refined answer failures now return sanitized structured output, preserve relevance/boundary receipts plus attempted answer usage, and skip support and vision stages.
  3. Legacy and --no-refine responses now include route, evidence status, confidence basis, warnings, embedding usage, answer usage, and sanitized failure receipts.
  4. Scene expansion now retains the full containing temporal event, including 10-20 caption plus 5-25 transcript and point-caption plus 10-15 transcript regressions.

Verification: 267 offline tests passed; source distribution and wheel built successfully; git diff --check passed. No live model calls were made.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent runtime re-review of commit 75125c3104278dd137e4960927e0afb047c81f90.

The two previously identified P2 defects are resolved:

  • Answer-provider failures retain accumulated Jev receipts and the failed answer request's usage snapshot in structured results. The same handling covers the refinement-outage fallback path.
  • Scene expansion adopts the containing temporal event's union bounds before the single-event early return. Regression cases cover an overlapping transcript wider than a caption and a point caption contained in a transcript.

Also verified that answer generation sends the configured chat_max_output_tokens limit and that legacy/--no-refine responses report answer and query-embedding usage. Failure responses sanitize provider details.

Validation at this exact commit: tests/test_ask_refinement.py, tests/test_provider.py, and tests/test_config.py — 59 passed. HEAD was checked before and after execution. GitHub CI reports success on Python 3.11, 3.12, and 3.13.

No open P1/P2 findings in this runtime review scope. Rejection handling, broad-summary isolation, successful-frame citation validation, opt-in fallbacks, and incomplete-usage behavior were included in the review. This review makes no claim about live model accuracy or measured cost savings; those require the separate reproduction receipts.

@seanphan

Copy link
Copy Markdown
Contributor Author

[P2] Natural-language questions can return no candidates before refinement runs

Confirmed at 75125c3104278dd137e4960927e0afb047c81f90 with an isolated offline database containing one caption, The truck was blue:

  • What color was the truck? → zero retrieval hits; refined_no_results; neither Jev nor the answer model is called.
  • What color was the truck → the same zero-hit result.
  • truck and "color" OR "truck" → the expected caption is retrieved.

search/rag.py:165–166 passes the full question into search/semantic.py:36, which forwards it directly to FTS MATCH. Question-mark punctuation survives db/repository.py:26–32 sanitization, causing the error fallback to use the full question as a LIKE substring. Without punctuation, implicit FTS AND still requires question words such as what and color to occur in the caption. This inherited retrieval behavior prevents the new refinement stage from running on ordinary supported questions.

Please add a generic natural-language candidate query path for av ask: bounded, safely quoted content terms combined for recall, with the original question retained for embedding/refinement/answering. Preserve the existing advanced FTS semantics of av search. Regression coverage should include ordinary questions with and without punctuation, video isolation, an unrelated/no-content query, and unchanged explicit FTS behavior. No answer-specific keyword or prompt rewriting is needed.

This finding extends the earlier review beyond the four fixes; it leaves one P2 open at this head.

@seanphan

Copy link
Copy Markdown
Contributor Author

Implemented the natural-language retrieval fix in 86ebf7c81bc3391213162a66641db42c8fd8ec70.

av ask now explicitly selects a generic candidate-query path: Unicode word tokens are quoted and OR-combined, with a 24-term cap applied after generic English question/function-word removal and deduplication. Punctuation cannot become FTS syntax. Empty-content questions return no candidates. The original question is still passed unchanged to embeddings, Jev, and answer synthesis. av search retains its existing advanced FTS behavior.

Offline regressions cover the generic truck question with and without punctuation, Unicode/punctuation, video isolation, stopword-only and unrelated questions, term limits, explicit FTS queries, original-question embeddings, and Jev rejection of broad candidates.

Validation before commit: 278 tests passed across the full offline suite; git diff --check passed. Targeted runtime suite: 70 passed. No live API calls were used. The earlier P2 reproduction is fixed; independent review of this new commit is pending.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent review of natural-language candidate retrieval at exact head 86ebf7c81bc3391213162a66641db42c8fd8ec70: no actionable P1/P2 findings in the four-file change.

Verified that av ask builds bounded, quoted OR candidates using generic Unicode lexical terms; punctuation and FTS operators cannot become query syntax. The original question remains the input to embeddings, Jev and answer synthesis. av search retains its existing explicit FTS semantics. Empty/function-word-only inputs return no candidates, video isolation remains in place, and the tests use generic fixtures without benchmark-specific vocabulary.

Independent validation: tests/test_natural_language_search.py passed all 11 tests on the reviewed checkout. Review was read-only with no model calls. This closes the natural-language retrieval delta only; subsequent provider-budget changes require review at their own exact head.

@seanphan

Copy link
Copy Markdown
Contributor Author

PASS — independent review of exact head bb22c76fe8f2da0b92c5ca9f637af2059fd69551, emphasizing 86ebf7c81bc3391213162a66641db42c8fd8ec70..bb22c76.

No actionable findings. api_token_limit_parameter is restricted to max_tokens or max_completion_tokens; the shared request helper emits exactly the selected field for frame captions, chunk captions, answers, summaries, and setup validation. Positive vision/chunk caps load from config and environment, flow through CLI-loaded config, and are preserved by the OAuth OpenAI fallback. I found no secret exposure, private implementation leakage, or behavioral/security regression in this delta.

Offline verification:

  • UV_PROJECT_ENVIRONMENT=/tmp/pixelml-av-pr2-review.jhzRbN/venv uv run --project /tmp/pixelml-av-pr2-review.jhzRbN/repo --frozen pytest -q tests/test_provider.py tests/test_config.py → 31 passed
  • same environment, pytest -q → 283 passed
  • env/CLI plus fake-client probes → caps 33/66/99/99, only max_completion_tokens; OAuth chunk cap 82, no max_tokens
  • git diff --check 86ebf7c81bc3391213162a66641db42c8fd8ec70..bb22c76fe8f2da0b92c5ca9f637af2059fd69551 → clean
  • checkout remained detached and clean; GitHub checks for Python 3.11/3.12/3.13 are successful

This is a review comment only; I did not approve or merge.

@seanphan

Copy link
Copy Markdown
Contributor Author

PASS — independent re-review of exact head 7e1a0ce, focused on 86b849d..7e1a0ce.

All four prior findings are resolved:

  1. Temporary branch permalinks and stale “unmerged” wording were replaced with repository-relative links; all changed targets exist.
  2. Receipt-level reservations now carry explicit lifecycle states. Known estimates remain $0.3913266; retained reservations are $1.11; known plus retained is $1.5013266, leaving $3.4986734 under the $5 cap. Historical ceilings total $3.98945825, including $2.87945825 released after metering.
  3. The three provider-cap variables are documented with accurate scope: primary OpenAI-compatible caption/chunk/answer/summarization calls only, excluding Jev, stronger inspection, bench, and sentinel; no upstream enforcement is claimed.
  4. Receipt sanitization wording now discloses the retained benchmark question/answer pair and why it is needed. Scans found no credentials, private paths/IPs, upload URIs, private routes, media, or transcript/caption corpora.

Validation at the exact detached head: 16 cost-model tests, 5 transcript-helper tests, 31 provider/config tests, notebook check, full 283-test suite, and sdist/wheel build all passed. GitHub CI is green on Python 3.11–3.13. No API calls, approval, merge, or publication occurred.

This is a review comment only; the PR remains draft.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent exact-head review — PASS

Reviewed PR #2 at e24845de7b58eb55062bbea4d5efe72e8c37d30b (base range 6a1cde2..e24845d), limited to the nine cookbook-file changes.

  • Restored-route receipt matches the recorded usage: one request, zero retries, no hidden fallback, exact model requested/returned, max_completion_tokens=32, 260 completion / 1168 prompt / 1152 cached / 0 reasoning tokens, complete usage, cap_enforced=false, cache-aware estimate $0.0009004, and source/frame hashes plus dimensions retained without raw media.
  • Cost model reconciles exactly: known total $0.3942579; reservation history $3.99945825 ($2.89945825 released after metering, $1.10 superseded, $0 retained); remaining headroom $4.6057421.
  • Documentation consistently states that the selected output-cap field was ignored, not enforced, and that no paid ingestion/query or completed AV comparison followed.
  • Offline validation passed: 16 cost-model tests, 5 transcript-helper tests, build_nb.py --check, and git diff --check.
  • Leakage review of the changed files found no credentials, upload identifiers, private endpoints/routes, raw source media, transcript/caption corpora, or unrelated private implementation material. The intentionally disclosed baseline benchmark question/answer pair was treated as authorized.

Provenance limitations: costs are cache-aware list-rate estimates, not provider invoices; infrastructure/storage/network costs remain unknown. Review did not re-run model or provider traffic.

@seanphan

Copy link
Copy Markdown
Contributor Author

STATE — fresh authorized live attempt recorded at dd9dfa2.\n\n- Transcript import completed (75 artifacts), but local media-probe preflight timed out before 300-frame extraction.\n- 0 captions, 0 provider requests, $0 new provider estimate; no retry and no Jev/Grok query against incomplete evidence.\n- Known cumulative list estimate remains $0.3942579 with $4.6057421 estimated headroom under the unchanged $5 cap.\n- Added sanitized receipt and updated cost ledger/notebook/tests; no media, corpus, credentials, private routes, or private implementation included.\n- Offline checks: 283 pytest tests passed; 16 cost-model tests passed; deterministic notebook current; uv build passed; git diff --check clean.\n\nPR remains draft/open/unmerged. No speed, cost, or quality comparison claim exists.

@seanphan

Copy link
Copy Markdown
Contributor Author

STATE — live Jev comparison arm closed by credential absence; no Jev request was made.

Done:

  • Confirmed no AV_TYPESAFE_API_KEY / TYPESAFE_API_KEY environment entry, credential-name match, or credential filename was present in the searched AV environment/config roots. No secret value was read or printed.
  • Preserved the completed ingestion database (300/300 caption requests, 75 transcript windows) and the completed Grok-only query receipt. Neither paid run was repeated.
  • Recorded the recovered query observer limitation: no Jev request ever ran with the incorrect Jev meter rates. Any future live Jev run must meter ordinary and list-token fields with the corrected $0.042/M input rate and retain raw usage.

Blocker:

  • A real Jev/TypeSafe credential. Because AV requires AV_TYPESAFE_API_KEY, the query route correctly fell back to legacy; substituting Grok for Jev would mislabel the result, so no completed Jev comparison is claimed.

Next:

  • Reconcile the ledger and public cookbook to the successful Grok-only query and this credential blocker, then rerun checks and request independent exact-head review. No merge/release/deploy/publication.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent exact-head review — PASS

Reviewed PR #2 at 59a7f7317329e4bba6e1b43cdf289a233e865a75, focused on dd9dfa2..59a7f73. No merge, approval, release, deployment, or publication action taken.

  • Cost arithmetic reconciles exactly: known token-derived list estimates total $0.9006585; retained reservations are $0. The reservation history records $5.99945825: $3.99945825 released after metering and $2.00 superseded, leaving $4.0993415 under the $5 estimate cap.
  • The 300-frame Grok ingestion receipt records 300/300 successful requests, complete aggregate usage, 300 persisted captions, 75 transcript windows, zero retries/fallbacks, a $0.5048525 cache-aware list estimate, and six responses above the requested 200-token advisory cap (maximum 235). The documentation does not claim that output caps were enforced.
  • The query receipt is correctly labeled Grok-only: one request, complete usage, a transcript citation, raw/unjudged evidence, and zero Jev requests because the required credential was absent. The blocked Jev receipt independently records zero requests and explicitly prevents substituting Grok for Jev.
  • Wording remains truthful: no completed Grok+Jev comparison, no all-in latency/cost comparison, and no aggregate quality-parity or savings claim is made from the single question.
  • New receipts and changed cookbook files passed the leakage review: no credentials, private endpoints/routes, upload IDs, source media, transcript/caption corpora, Sean X draft material, or private Composer implementation.
  • Checks rerun read-only in a detached temporary checkout: cookbook/cost-model/test_model.py passed 16 tests and 11 subtests; build_nb.py --check reported notebook source and outputs current; git diff --check dd9dfa2..59a7f73 was clean.

Provenance: dollar figures are cache-aware public list-rate estimates, not invoices; provider/proxy billing and infrastructure/storage/network costs remain unknown. Full tests/build were reported as passing by the operator and were not repeated in this independent checkout.

@seanphan

Copy link
Copy Markdown
Contributor Author

Exact-head review — PASS

Head: dad9dcf
Scope: docs-only delta from reviewed 59a7f73 (10 cookbook files).

Offline validation in a detached read-only checkout:

  • Ancestry: 59a7f73 is an ancestor of dad9dcf; checkout remained clean.
  • pytest cookbook/cost-model/test_model.py -q: 16 passed, 12 subtests passed.
  • pytest cookbook/jev-refined-ask/test_transcribe_gemini.py -q: 5 passed.
  • cookbook/cost-model/build_nb.py --check: notebook source and outputs current.
  • model.py scenario.receipts.json: known experiment total 1.04309935 against 5 cap; retained 0; remaining 3.95690065; historical reservations 6.19945825.
  • git diff --check 59a7f73..dad9dcf: clean.
  • Sanitized scan over all ten changed files and all added lines: no credential values, private endpoints/routes, upload IDs/URIs, source media, transcript/caption corpus, Sean X draft, or private Composer implementation. The benchmark question/answer pair remains intentionally disclosed for reproducibility.

New receipt checks:

  • Stage costs 0.079296 + 0.061488 + 0.00165685 = 0.14244085.
  • Known cumulative list-rate estimate is 1.04309935 of the 5 cap.
  • Requests are exactly 2 Jev (relevance, support) + 1 Grok answer; automatic retries and fallbacks are both 0.
  • Usage totals reconcile to 4,533 input, 236 output, 128 cached, and complete per-stage usage is recorded.
  • Documentation consistently describes one completed Jev-refined question, no aggregate quality/speed/all-in-cost claim, unknown billed and infrastructure costs, and non-enforced output caps.

Provenance limitation: this is an offline exact-head audit of recorded receipts and docs; external provider responses were not replayed or billed.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent evidence review of dad9dcf5b53baf080f1242b4159dc405ab06fbcd: changes requested before publishing cost claims.

[P1] Correct the Jev input-rate unit before deriving query and cumulative totals. In the refined-query receipt, relevance has 1,888 input tokens and an estimate of $0.079296; support has 1,464 input tokens and $0.061488. Those estimates apply $42 / 1,000,000 tokens. TypeSafe's current public pricing says “$42 Per Billion input tokens”, equivalent to $0.042 / 1,000,000. The two input estimates should be $0.000079296 and $0.000061488, totaling $0.000140784 (exactly 1,000× below the checked-in judge subtotal).

The Grok answer's cache-aware list estimate of $0.00165685 reconciles with 1,181 input tokens, 128 cached input tokens and 126 output tokens. Adding that answer estimate to the corrected Jev input estimate produces $0.001797634 for this recorded query, assuming no separate Jev output charge. Actual provider/proxy billing remains unknown; verify and document the output-rate basis explicitly.

Propagate the unit correction to scenario.receipts.json, the receipt's cumulative total, the cost-model README, and generated notebook/output. Preserve the raw token counts and failed-attempt accounting. Add a regression that uses 1 billion Jev input tokens and expects $42, so arithmetic-only tests cannot miss the rate-unit mistake. No new paid inference is necessary to fix these derived estimates.

@seanphan

Copy link
Copy Markdown
Contributor Author

Implemented the Jev pricing correction at bcf6c12439bea65f2eaea10f5449c644e8f4e765. The official TypeSafe documentation states $42 per billion input tokens ($0.042 per million) with free output.

The unchanged 1,888 relevance and 1,464 support input tokens produce $0.000140784 in Jev list-price estimates. Adding the unchanged $0.00165685 Grok answer estimate gives $0.001797634 for the query's model calls. The known experiment estimate is $0.902456134, with $4.097543866 of estimate headroom under the $5 cap. These are not verified billed dollars.

The receipt retains the original derived values and a dated correction. Measured usage, answer content and timings are unchanged. The regression derives the price from the published billion-token unit, includes free output, and checks both stages against the query and cumulative ledger. The notebook was regenerated. Related provenance now explicitly states that the answer is a normalized factual excerpt, Gemini implicit cache usage is unknown, and the ingestion timer excludes earlier frame extraction and external ASR.

Validation: 283 AV tests and 18 cost-accounting tests passed; notebook consistency and diff checks passed. No provider requests were made. Independent review of this exact commit has been requested.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent review at exact head bcf6c12439bea65f2eaea10f5449c644e8f4e765: the P1 pricing-unit finding is resolved. PASS for the bounded pricing, receipt-provenance, and escalation-path review. No new blocker found in this scope.

The correction now derives TypeSafe's published $42 per billion input tokens = $0.042 per million, with free output, from the official model documentation. The query receipt reconciles:

Metered stage List-price estimate
Jev relevance: 1,888 input tokens $0.000079296
Jev answer support: 1,464 input tokens $0.000061488
Grok answer, including reported cache usage $0.00165685
Completed AV query $0.001797634

The original $0.14244085 query estimate is retained as correction history. The $0.140643216 reduction reconciles the experiment ledger and generated notebook to $0.902456134 in known list-price estimates, with $4.097543866 estimate headroom under the $5 cap. Raw measured usage and timings were preserved; the four correction-recorded raw receipt checksums still match. Actual billing, infrastructure allocation, and the unmetered failed request remain unknown.

The completed one-time preparation is $0.5048525 captions + $0.0685709 ASR = $0.5734234. The first AV question including that preparation is $0.575221034, versus the native Gemini receipt's $0.308076 full-rate estimate. These are model-call estimates, not complete service costs or invoices. The receipt/notebook now correctly leave Gemini's unreported implicit cache usage unknown.

The answer is correctly labeled a normalized factual excerpt, not the complete verbatim returned response. Its core 12.2% manufacturing-growth answer agrees with the recorded transcript evidence and native response; the AV citation is the coarse 19:00–20:00 transcript window. This does not validate every omitted timestamp/visual statement, establish representative accuracy, or turn the Jev support score into a correctness probability.

The reviewed runtime remains unchanged by this correction. In rag.py, no hits and successful evidence rejection stop without a supported answer; a Jev outage is explicitly marked raw/unjudged. Weak or unknown support can use the explicitly configured inspector, and its result is support-judged again. Inspection has bounded sampled-still budgets; it does not send native video/audio or establish events between frames. No stronger inspection occurred in the live query, so its quality benefit remains unmeasured.

Independent offline checks on this exact head: 18 cost-accounting tests plus 12 subtests passed; 37 focused AV refinement/ingestion tests passed; deterministic notebook check and git diff --check passed. No provider calls were made. The 3.7385s AV and 50.4044s Gemini timings cover one query after preparation; the AV ingestion timer excludes earlier frame extraction and ASR. No aggregate accuracy-parity, universal savings, or complete-cost break-even claim is approved by this review.

@seanphan

Copy link
Copy Markdown
Contributor Author

Independent exact-head residual review — PASS

Head: bcf6c12439bea65f2eaea10f5449c644e8f4e765. No merge, approval, release, or publication action taken.

  • The dad9dcf..bcf6c12 delta is limited to cost-model/receipt/documentation updates; no runtime AV code changed. The prior runtime audit therefore remains applicable: genuine Jev rejection stays empty, provider outage is explicitly raw/unjudged, support is a separate decision, stronger inspection is opt-in and bounded, natural-language retrieval and transcript validation are covered, and usage/cap receipts are preserved.
  • The Jev correction is internally consistent: $42/B input with free output becomes $0.042/M; the two Jev stages total $0.000140784, the recorded Grok answer is $0.00165685, and the query total is $0.001797634. The ledger/notebook total is $0.902456134 with no retained reservation. The regression derives the unit explicitly and rejects unknown units.
  • The evidence remains bounded to one completed question. The receipts preserve measured token/timing data, distinguish list-rate estimates from billing, leave implicit Gemini cache usage and unpriced infrastructure unknown, and make no universal savings, parity, or aggregate quality claim.
  • I found no new credential, private endpoint, proprietary prompt/algorithm, payload corpus, source media, or infrastructure disclosure in this delta. No provider requests were made and no redundant full-suite rerun was needed.

PASS for the residual code-correctness and disclosure-boundary review at this exact head. CI 3.11/3.12/3.13 is green.

@seanphan
seanphan marked this pull request as ready for review September 19, 2026 18:03
@seanphan
seanphan merged commit 788c046 into main Sep 19, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant