Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,7 @@ av ingest /folder/ # batch directory
av ingest video.mp4 --force # re-ingest
av ingest "https://youtu.be/..." # YouTube URL
av ingest video.mp4 --dense-vision # structured dense captions
av ingest video.mp4 --transcript-json transcript.json # validated external transcript; skips built-in ASR
```
```json
{"status": "complete", "video_id": "uuid", "filename": "video.mp4", "duration_sec": 120.5, "artifacts_count": 42, "elapsed_sec": 15.3}
Expand All @@ -127,6 +128,12 @@ On partial failure: `{"status": "complete_with_warnings", ..., "warnings": ["Tra
{"answer": "...", "citations": [{"video_id": "uuid", "start_sec": 120.0, "source_type": "transcript", "text": "...", "score": 0.91}], "confidence": 0.85}
```

With `AV_TYPESAFE_API_KEY` (or `TYPESAFE_API_KEY`) configured, `av ask`
automatically runs Jev source relevance, configurable bounded scene grouping, answer
synthesis, and a separate answer-support Noul. `--no-refine` preserves legacy RAG
for one request. Refined responses retain answer/citations/confidence and add route,
evidence, refinement, warning, inspected-window, and stage-usage metadata.

### `av list` / `av info <id>` / `av transcript <id>` / `av export` / `av open <id>`
See `av <command> --help` for details.

Expand Down Expand Up @@ -180,12 +187,28 @@ When a capability is unavailable (e.g. Anthropic has no Whisper), the pipeline s
|----------|---------|-------------|
| `AV_API_KEY` | (none) | API key (overrides config.json) |
| `AV_API_BASE_URL` | `https://api.openai.com/v1` | API endpoint |
| `AV_API_TIMEOUT_SEC` | `120` | Per-attempt API timeout |
| `AV_API_MAX_RETRIES` | `1` | Explicit retry count; SDK retries stay disabled |
| `AV_ALLOW_OAUTH_FALLBACK` | `false` | Allow reading local OAuth caches only when explicitly enabled |
| `AV_ALLOW_CODEX_FALLBACK` | `false` | Allow spawning Codex for caption fallback only when explicitly enabled |
| `AV_PROVIDER` | (none) | Provider name |
| `AV_TRANSCRIBE_MODEL` | `whisper-1` | Transcription model |
| `AV_VISION_MODEL` | `gpt-4-1` | Vision/caption model |
| `AV_EMBED_MODEL` | `text-embedding-3-small` | Embedding model |
| `AV_CHAT_MODEL` | `gpt-4-1` | Chat/RAG model |
| `AV_CHAT_MAX_OUTPUT_TOKENS` | `1024` | Positive output-token cap for each answer response |
| `AV_DB_PATH` | `~/.config/av/av.db` | Database location |
| `AV_TYPESAFE_API_KEY` | (none) | Jev/System One key; enables ask refinement by default |
| `AV_TYPESAFE_ENDPOINT` | `https://api.typesafe.ai/v1/systemone` | Explicit System One endpoint |
| `AV_TYPESAFE_MODEL` | `jev-latest` | System One model |
| `AV_REFINE_RELEVANCE_MIN` | `0.5` | Minimum source-relevance Noul probability |
| `AV_REFINE_SUPPORT_MIN` | `0.5` | Minimum answer-support Noul probability |
| `AV_REFINE_MAX_SCENES` | `8` | Maximum merged scenes sent to synthesis |
| `AV_REFINE_BATCH_SIZE` | `10` | Sources per System One relevance request |
| `AV_REFINE_CONTEXT_EVENTS` | `3` | Maximum temporal events on each side of a hit |
| `AV_STRONG_VISION_API_BASE_URL` | (none) | Explicit OpenAI-compatible sampled-frame endpoint |
| `AV_STRONG_VISION_API_KEY` | (none) | Key for the stronger sampled-frame endpoint |
| `AV_STRONG_VISION_MODEL` | (none) | Explicit stronger vision model; never selected implicitly |
| `DEEPSEEK_API_KEY` | (none) | Key for a self-hosted DeepSeek-V4.1-Flash server |
| `SGLANG_API_KEY` | (none) | Alias for the same, matching SGLang's own naming |

Expand Down
102 changes: 101 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,63 @@ av search "person with red bag"
av ask "what happened at 2:30?"
```

### Refined `av ask` (optional)

Configure a TypeSafe System One key to make Jev refinement automatic for `av ask`:

```bash
export AV_TYPESAFE_API_KEY="..." # TYPESAFE_API_KEY is also accepted
av ask "when does the person enter the room?"
av ask "when does the person enter the room?" --no-refine # legacy RAG for this call
```

The refined path uses Jev only for typed decisions: it filters source relevance,
uses a configurable bounded temporal neighborhood to form local scenes, merges
overlapping same-video scenes, and ranks them by `relevance probability × retrieval
score`. A separate Jev Noul checks whether the answer is supported; source relevance
is not treated as answer correctness.

If Jev is unavailable, `av ask` visibly warns and falls back to raw retrieval. If
Jev validly rejects every hit, the result is empty instead of restoring rejected
hits. Refined JSON includes `route`, `evidence_status`, `refinement`, `warnings`,
`inspected_windows`, and per-stage token usage when providers report it. Unknown
usage remains `null`.

Missing, malformed, or out-of-range System One probabilities are treated as a
refinement failure: `av` reports the fallback and does not invent a confidence.

See the [AV ask refinement cookbook](cookbook/README.md) for the reproducible
recipe, offline cost arithmetic, and sanitized receipt provenance.

FTS5 remains the first retrieval stage. An unscoped query with zero FTS matches does
not scan the video archive or invoke sampled-frame inspection.

### API-only reproducible ingest

Use an explicit OpenAI-compatible endpoint/key and keep both fallback flags disabled:

```bash
export AV_API_BASE_URL="https://your-provider.example/v1"
export AV_API_KEY="..."
export AV_VISION_MODEL="your-cheap-vision-model"
export AV_ALLOW_OAUTH_FALLBACK="false"
export AV_ALLOW_CODEX_FALLBACK="false"

av ingest video.mp4 --dense-vision --max-frames 120 --no-embed
```

To import a timestamped transcript produced by a separate public ASR script, pass a
validated sidecar instead of running built-in ASR:

```bash
av ingest video.mp4 --dense-vision --transcript-json transcript.json
```

The sidecar contains `segments` with `start_sec`, `end_sec`, and `text`, plus optional
`model` and public `provenance`. It is validated against the probed video duration
before database changes or API calls. AV records the import as local work with zero
provider requests; external ASR usage or cost is not attributed to this ingest.

### Surveillance Detection

```bash
Expand Down Expand Up @@ -242,17 +299,60 @@ Env vars always override config.json:
```bash
export AV_API_KEY="sk-..."
export AV_API_BASE_URL="https://api.openai.com/v1" # or any OpenAI-compatible endpoint
export AV_API_TIMEOUT_SEC="120"
export AV_API_MAX_RETRIES="1"
export AV_API_TOKEN_LIMIT_PARAMETER="max_tokens" # or max_completion_tokens when required
export AV_ALLOW_OAUTH_FALLBACK="false" # never read local auth caches unless explicitly enabled
export AV_ALLOW_CODEX_FALLBACK="false" # never spawn Codex unless explicitly enabled
export AV_TRANSCRIBE_MODEL="whisper"
export AV_VISION_MODEL="gpt-4-1"
export AV_VISION_MAX_OUTPUT_TOKENS="200" # single-frame caption response
export AV_VISION_CHUNK_MAX_OUTPUT_TOKENS="500" # multi-frame chunk caption response
export AV_EMBED_MODEL="text-embedding-3-small"
export AV_CHAT_MODEL="gpt-4-1"
export AV_CHAT_MAX_OUTPUT_TOKENS="1024" # positive cap for each answer response

# Optional Jev/System One refinement (automatic when a key is present)
export AV_TYPESAFE_API_KEY="..." # TYPESAFE_API_KEY also works
export AV_TYPESAFE_ENDPOINT="https://api.typesafe.ai/v1/systemone"
export AV_TYPESAFE_MODEL="jev-latest"
export AV_REFINE_RELEVANCE_MIN="0.5"
export AV_REFINE_SUPPORT_MIN="0.5"
export AV_REFINE_MAX_SCENES="8"
export AV_REFINE_BATCH_SIZE="10"
export AV_REFINE_CONTEXT_EVENTS="3"

# Optional bounded sampled-frame fallback after an unsupported answer
export AV_STRONG_VISION_API_BASE_URL="https://your-explicit-endpoint.example/v1"
export AV_STRONG_VISION_API_KEY="..."
export AV_STRONG_VISION_MODEL="your-explicit-model"
export AV_INSPECTION_MAX_WINDOWS="2"
export AV_INSPECTION_MAX_SECONDS="120"
export AV_INSPECTION_MAX_FRAMES="12"
export AV_INSPECTION_MAX_ATTEMPTS="1"
export AV_INSPECTION_DENSE_PASS="false"

# Self-hosted DeepSeek-V4.1-Flash via SGLang
export AV_PROVIDER="deepseek"
export AV_API_BASE_URL="http://your-sglang-host:30000/v1"
export DEEPSEEK_API_KEY="..." # only if your server requires one
```

`AV_API_TOKEN_LIMIT_PARAMETER` selects the request field sent by AV's primary
OpenAI-compatible chat provider. The two vision limits apply to its single-frame
and multi-frame caption calls; `AV_CHAT_MAX_OUTPUT_TOKENS` applies to its ordinary
answer and summarization calls. These settings do not configure Jev/System One,
stronger sampled-frame inspection, `av bench`, or `av sentinel`. They also do not
prove that an upstream provider accepts or enforces the requested cap; verify the
returned usage and finish reason for the exact endpoint and model.

API requests use the configured timeout and explicit retry limit. Ingestion JSON
includes `stage_usage` for transcription, captioning, caption summarization, and
embeddings, plus the effective frame/request settings. Request failures are counted;
token totals become `null` with a completeness flag when any provider omits usage.
Ask JSON likewise reports the effective chat model/output cap and per-stage usage.
No dollar total is inferred.

## Requirements

- Python 3.11+
Expand All @@ -267,7 +367,7 @@ export DEEPSEEK_API_KEY="..." # only if your server requires one
| `av config show` | Show current configuration |
| `av ingest <path>` | Ingest video file(s) into the index |
| `av search <query>` | Full-text + semantic search |
| `av ask <question>` | RAG Q&A with citations |
| `av ask <question>` | RAG Q&A; automatically refines with Jev when configured |
| `av list` | List all indexed videos |
| `av info <video_id>` | Detailed video metadata |
| `av transcript <id>` | Output transcript (VTT/SRT/text) |
Expand Down
21 changes: 21 additions & 0 deletions cookbook/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# AV cookbook

Runnable recipes for the open-source **av** CLI live here alongside the code.

| Recipe | What it demonstrates | Evidence status |
|---|---|---|
| [Cost model](cost-model/README.md) | Separate tokens, estimates, unknown costs, ingestion, queries, failures, reservations, and cap accounting | Completed one-question comparison; no aggregate parity claim |
| [Jev-refined ask](jev-refined-ask/README.md) | Build a source-bound transcript sidecar, retrieve indexed moments, refine evidence, answer, and check support | Runnable recipe; no speed, cost, or quality parity claim |
| [Sanitized receipts](receipts/README.md) | Completed ASR/baseline, caption smoke/abort, incompatible cap probes, Grok ingestion, and Jev-refined query | No media, transcript/caption corpus, credentials, upload URIs, or private routes; baseline question and returned answer/rationale retained |

The [original public cost notebook](https://github.com/PixelML/cookbook/tree/main/agentic-video/cost-model)
remains available at its existing URL. Its Composer demonstrations are historical
context, not measurements of this CLI. New AV reproduction results belong in
this cookbook with their own media, model, configuration, and cost provenance.

The current evidence includes a completed Gemini 3.8 direct-video baseline, a
completed 300/300 Grok caption ingestion with 75 transcript windows, one Grok-only
answer, and one later Jev-refined answer for the same question. The selected route
ignored output caps in both 32-token probes, and six of the 300 caption responses
exceeded the requested 200-token advisory cap. The paired result is one question;
do not infer aggregate speed, cost, or quality parity from it.
Loading
Loading