diff --git a/.gitignore b/.gitignore index e599c84..1bed3f5 100644 --- a/.gitignore +++ b/.gitignore @@ -10,6 +10,9 @@ node_modules/ # TrueForge local data (SQLite lives in the OS app data dir; this is just a run dir) .trueforge/ +# ffmpeg sandbox scratch space: media the agent works on, never source +packages/ffmpeg-sandbox/workspace/ + # python bytecode from the merge-gate parser and its tests __pycache__/ *.pyc diff --git a/README.md b/README.md index f8d0f28..5487fdb 100644 --- a/README.md +++ b/README.md @@ -1,10 +1,28 @@ +
+ # EditAI -**An AI video editor you talk to.** Say "remove the silences" or "caption every clip", and an agent -makes the edit on your real timeline: it reads the project, decides which cuts to make, and applies -them. Anything destructive stops and asks you first. +**An AI harness for video editing. You describe the edit; an agent makes it on your real timeline.** + +[**Live site**](https://editai-agent.vercel.app) · +[Quick start](#getting-started) · +[The 19 tools](#the-19-timeline-tools) · +[Add your own MCP](#reaching-past-the-timeline) · +[Review evidence](#code-review-evidence-qodo) [![License: MIT](https://img.shields.io/badge/license-MIT-green?style=flat)](LICENSE) +[![Built on TrueForge](https://img.shields.io/badge/harness-TrueForge-7c5cff?style=flat)](https://trueforge.dev) +[![MCP](https://img.shields.io/badge/protocol-MCP-7c5cff?style=flat)](https://modelcontextprotocol.io) +[![Tests](https://img.shields.io/badge/tests-71%20passing-3aa39b?style=flat)](#tests) +[![Reviewed by Qodo](https://img.shields.io/badge/reviewed%20by-Qodo-e0a63b?style=flat)](#code-review-evidence-qodo) + +
+ +--- + +Say "remove the silences" or "caption every clip", and an agent makes the edit on your real +timeline: it reads the project, decides which cuts to make, and applies them. Anything destructive +stops and asks you first. The tedious parts of editing are the ones a machine should do. Cutting dead air out of a twenty-minute take is thirty minutes of scrubbing; here it is one sentence and one approval click. @@ -12,50 +30,94 @@ Captioning every clip means transcribing each one by hand; here a sub-agent hand parallel and the captions land on their own track, timed. The agent does the mechanical work, and you keep the decisions: it proposes, you approve, and every edit is undoable. -### What it can do today +**Three layers, one conversation.** Depending on what you ask for, the same agent drives your own +editing UI, the editor in this repo, or the ffmpeg encode underneath it. You do not choose the +layer; the request does. -| Ask for this | What the agent does | -| --- | --- | -| *"Remove the silences"* | Finds every silent range on the voice track and ripple-deletes it across all tracks, closing the gaps. Verified: 24s → 21.1s, exactly the 2.9s of silence. | -| *"Caption every video clip"* | Fans out one sub-agent per clip to transcribe in parallel, merges the results, and lays timed captions on a new track. | -| *"Cut the intro to 3 seconds"* | Trims the clip, keeping the media in sync by moving its source offset. | -| *"Duck the music under the voiceover"* | Sets clip volume. | -| *"Split this at 15 seconds"* | Cuts a clip in two, both halves still frame-accurate. | -| *"Export it at 1080p"* | Renders, after you approve. | +--- + +## Watch it work -### The agent tools +The one operation worth showing end to end, because the numbers are checkable: -The timeline is exposed to the agent as **16 [MCP](https://modelcontextprotocol.io) tools**, not as -a prompt describing a timeline. The agent calls real functions against real state: +``` +you ▸ remove the silences -- **Read:** `get_project`, `transcribe_clip`, `find_silences`, `detect_beats`, `list_changes` -- **Write:** `split_clip`, `trim_clip`, `move_clip`, `set_volume`, `add_text`, `add_captions`, `undo` -- **Destructive:** `delete_clip`, `ripple_delete`, `remove_silences` -- **Gated:** `export_project` +agent → find_silences(min_duration: 0.5, track_id: "a1") + ← 3 ranges: 3.2-4.1, 9.6-10.4, 16.8-18.0 (2.9s total) -Every tool validates its input and returns errors to the model as data, so a stale clip id becomes -a correction the agent recovers from rather than a failed turn. The destructive four are published -with MCP's `destructiveHint` annotation, which is what makes the harness stop and ask you before -they run. + "Three silent ranges, 2.9s in all. This will cut every track. Remove them?" -Because it is MCP, the same agent can reach anything else that speaks MCP: web search, your issue -tracker, an internal API you wrap yourself. Connectors attach by name and authorize in chat. +you ▸ [Approve] + +agent → remove_silences(min_duration: 0.5) + ← duration 24.0s → 21.1s, every track cut, gaps closed + + "Cut 2.9s of dead air across every track. Say undo to put it back." +``` + +24.0 down to 21.1 is exactly the 2.9 seconds of silence, with picture, sound and captions still in +sync. There is an animated version of this on the [live site](https://editai-agent.vercel.app). + +## What you can ask for + +| Ask | What the agent actually does | Tools | +| --- | --- | --- | +| *"Remove the silences"* | Finds every silent range on the voice track, tells you what will go, then ripple-deletes it across all tracks, closing the gaps. Verified: 24s → 21.1s. | `find_silences` → `remove_silences` | +| *"Caption every video clip"* | Fans out one sub-agent per clip to transcribe in parallel, merges the results (including sentences that straddle a cut), and lays timed captions on their own track. | `transcribe_clip` ×N → `add_captions` | +| *"Cut this to the beat"* | Reads the tempo off the music track and splits on the beat grid. Both halves stay frame-accurate. | `detect_beats` → `split_clip` | +| *"Trim the intro to 3 seconds"* | Trims the clip and moves its source offset by the same amount, so the picture does not jump. | `trim_clip` | +| *"Duck the music under the voiceover"* | Sets clip volume where the voice track is speaking and restores it where it is not. | `get_project` → `set_volume` | +| *"Grade it warmer and add grain"* | Writes the ffmpeg filter graph, runs it in the media sandbox, then probes the output to confirm it matches the intent. | `probe_media` → `run_ffmpeg` | +| *"Kill the room tone"* | Denoises the voice track in the sandbox, leaving the original file untouched beside it. | `run_ffmpeg` (`afftdn`) | +| *"Put their logo in the corner"* | Searches the live web through Bright Data, scrapes the asset, and brings it into the project. | `search_engine` → `scrape_as_markdown` | +| *"Export it at 1080p"* | Renders, after you approve. | `export_project` | + +Motion graphics, transitions and animation work the same way: either as an ffmpeg filter graph in +the media sandbox, or by attaching an MCP server that specialises in them. See +[Reaching past the timeline](#reaching-past-the-timeline). ## How it works -EditAI is built on [TrueForge](https://trueforge.dev), an open-source agent harness. The harness -runs the agent loop; EditAI supplies the domain. +EditAI is built on [TrueForge](https://trueforge.dev), TrueFoundry's open-source agent harness. The +harness runs the agent loop; EditAI supplies the domain. ``` browser harness domain ┌──────────────┐ HTTP ┌──────────────┐ MCP ┌──────────────────┐ │ editor UI │◄─────────►│ TrueForge │◄────────►│ @editai/agent │ -│ (apps/web) │ + SSE │ agent loop │ │ 16 timeline │ -│ │ │ approvals │ │ tools │ -│ timeline ◄──┼───────────┼──────────────┼── SSE ───┤ project store │ -└──────────────┘ └──────────────┘ └──────────────────┘ +│ (apps/web) │ + SSE │ agent loop │ │ 19 timeline │ +│ decode │ │ approvals │ │ tools │ +│ composite │ │ │ │ project store │ +│ encode │ │ │ │ media on disk │ +│ timeline ◄──┼───────────┼──────────────┼── SSE ───┤ render queue │ +└──────┬───────┘ └──────┬───────┘ └────────┬─────────┘ + │ │ MCP │ + │ media bytes + rendered mp4 (HTTP) │ + └───────────────────────────────────────────────────────┘ + │ + ┌─────────────┼──────────────┐ + ▼ ▼ ▼ + ┌──────────────┐ ┌─────────┐ ┌──────────────┐ + │ ffmpeg │ │ bright- │ │ anything │ + │ sandbox │ │ data │ │ else that │ + │ (container) │ │ (web) │ │ speaks MCP │ + └──────────────┘ └─────────┘ └──────────────┘ ``` +**The editor is the renderer.** Decoding and encoding happen in the browser through +[WebCodecs](https://developer.mozilla.org/en-US/docs/Web/API/WebCodecs_API), wrapped by +[mediabunny](https://mediabunny.dev): the agent queues a render, the editor claims it, composites +every frame onto a canvas, muxes it, and posts the file back. The agent then sees a real path and a +real byte count, which is what lets it check its own work. + +The engine follows [OpenCut](https://github.com/opencut-app/opencut-classic), whose renderer this is +ported from: the same frame cache (a forward iterator with a prefetched next frame, falling back to a +real seek only when the target is behind the decoder or too far ahead of it) and the same +`Output`/`CanvasSource`/`AudioBufferSource` muxing. The one deliberate departure is compositing in +Canvas2D rather than OpenCut's wgpu compositor: EditAI stacks video, text and audio with no effects +or masks, which 2D covers exactly, and it drops a wasm dependency. Effects would need the real thing. + The editor never calls a model. It creates a session, streams turn events, renders tool calls and sub-agent threads, and answers the harness when it pauses. The timeline redraws from the agent server's event stream, so an edit the agent makes shows up in the UI as it happens. @@ -68,49 +130,240 @@ server's event stream, so an edit the agent makes shows up in the UI as it happe | Sessions that survive a reload | `apps/web/src/components/editor/use-assistant.ts` replays turns and re-attaches to a running one | | Live web research | `bright-data` connector: `search_engine` and `scrape_as_markdown`, attached deferred so it costs no context until a task needs it (see [docs/brightdata.md](docs/brightdata.md)) | | Any model provider | `apps/agent/scripts/setup.ts` registers whichever API keys are present, including any OpenAI-compatible endpoint | -| Sandboxed execution | Two layers: the harness sandbox (Daytona, configured automatically when `DAYTONA_API_KEY` is set) for general code, and `packages/ffmpeg-sandbox` for media work | +| Sandboxed execution | Two layers: the harness sandbox (Daytona) for general code, and `packages/ffmpeg-sandbox` for media work | +| Domain know-how | `skills/video-editing/SKILL.md`, loaded by the agent whenever a task touches ffmpeg | + +## The 19 timeline tools + +The timeline is exposed to the agent as **19 [MCP](https://modelcontextprotocol.io) tools**, not as +a prompt describing a timeline. The agent calls real functions against real state. + +| Tool | Kind | Arguments | What it does | +| --- | --- | --- | --- | +| `get_project` | read | | Tracks, clips, media metadata, exports. Call it first; ids change after edits. | +| `list_media` | read | | Imported files with their measured duration, resolution and frame rate, and whether the bytes are on disk. | +| `list_changes` | read | `limit` | Recent edits, oldest first. | +| `transcribe_clip` | read | `clip_id` | Speech inside one clip, as timed segments in timeline seconds. | +| `find_silences` | read | `min_duration`, `track_id` | Silent ranges measured from the decoded audio. Preview only. | +| `detect_beats` | read | `track_id` | Beat timestamps from the tempo estimated off the track's onsets. | +| `split_clip` | write | `clip_id`, `at` | Cuts a clip in two. Returns both halves. | +| `trim_clip` | write | `clip_id`, `start?`, `end?` | New in/out points, keeping media in sync via `sourceOffset`. | +| `move_clip` | write | `clip_id`, `start?`, `track_id?` | New start time, or another track of the same kind. | +| `set_volume` | write | `clip_id`, `volume` | Clip volume, 0 to 100. | +| `add_clip` | write | `name`, `track_id`, `start`, `duration?`, `source_offset?` | Places imported media on a video or audio track. | +| `add_text` | write | `text`, `start`, `duration`, `track_id` | A title or caption on a text track. | +| `add_captions` | write | `segments[]`, `track_label` | Timed captions on the captions track, creating it if needed. | +| `undo` | write | | Reverts the most recent change. | +| `delete_clip` | **destructive** | `clip_id` | Removes a clip, leaving a gap. | +| `ripple_delete` | **destructive** | `start`, `end` | Removes a range from every track and closes the gap. | +| `remove_silences` | **destructive** | `min_duration`, `track_id` | Ripple-deletes every silence over the threshold. | +| `export_project` | **approval** | `format`, `resolution` | Queues a real render. Refuses if any clip's media is missing. | +| `get_export` | read | `export_id` | Render status: pending, rendering with progress, done with the file and its byte size, or failed with the error. | + +Every tool validates its input with zod and returns errors to the model **as data**, so a stale clip +id becomes a correction the agent recovers from rather than a failed turn. + +The four gated tools are published with MCP's `destructiveHint` annotation, and the agent declares +`require_approval_for_tools: ["@destructive", "export_project"]`. That turns the annotation into a +pause: the harness stops the turn, the editor shows the tool and its arguments, and the run only +continues once a person allows or denies it. + +Three semantics worth knowing before you read the code: + +- **`sourceOffset` keeps media in sync.** Trimming a clip's start moves its offset into the source + file by the same amount, so the picture does not jump. +- **Ripple delete is the interesting operation.** Removing a range cuts every track, splits any clip + straddling the range, and shifts everything after it left. `remove_silences` applies it once per + silence, from the end backwards, so earlier ranges stay valid. +- **Captions merge across clip boundaries.** Fanning captioning out per clip means a sentence + spanning a cut is reported twice, clamped to each side; `add_captions` merges those back into one. + +See [apps/agent/README.md](apps/agent/README.md) for the full tool reference and timeline semantics. + +## Reaching past the timeline + +Because everything is MCP, the same agent can reach anything else that speaks MCP: web search, a +motion-graphics server, your issue tracker, an internal API you wrap yourself. **Adding a capability +is a name in a list and a restart, not a release.** + +```bash +# keyless, the default +EDITAI_CONNECTORS=exa bun run setup + +# header auth +BRIGHT_DATA_MCP_HEADER="Authorization: Bearer " \ + EDITAI_CONNECTORS=exa,bright-data bun run setup + +# OAuth: dynamic client registration, nothing to configure here. +# The first time the agent reaches for it, the turn pauses with an authorize +# URL and the editor shows a Connect button. Verified against Linear. +EDITAI_CONNECTORS=exa,linear bun run setup + +# your own server, attached the same way +EDITAI_CONNECTORS=exa,bright-data,motion-graphics bun run setup +``` + +Extra connectors attach **read-only and deferred**, so a connector you rarely use costs nothing in +context until the agent actually reaches for it. The tools go live on the agent's next turn, in the +same conversation. + +### Bright Data + +**Bright Data** is the one wired up and verified: `search_engine`, `search_engine_batch`, +`scrape_as_markdown`, `scrape_batch` and `ask_brightdata_assistant`, reported by the harness as +`auth_status: authenticated`. + +It attaches by name like any catalog connector, `EDITAI_CONNECTORS=exa,bright-data`, with its token +in `BRIGHT_DATA_MCP_HEADER`. With that variable unset, setup **skips** the connector rather than +registering it with an empty header, because a registered connector whose every call fails at run +time is worse than an absent one. + +What it is actually for is research that changes an edit, not decoration beside it: + +- **Naming and ordering clips.** Before titling, the agent researches how comparable short-form + videos are titled and applies the structures it finds. On Startup Boston Week footage it scraped + conference short-form titles, extracted four recurring shapes (number + highlight promise, + compressed recap, access/FOMO, question hook), and titled the clips to match. +- **Fetching real assets.** "Put their logo in the corner" gets the current logo rather than a + model's memory of one. + +Two properties worth stating, because they are what make live web data safe to act on: + +- **Scraped text is evidence, never instruction.** The harness wraps each payload in a notice + marking it untrusted external data, so a page cannot talk the agent into a tool call. +- **A layout change costs quality, not correctness.** `scrape_as_markdown` returns rendered text + rather than a selector path, and the agent treats an empty or truncated scrape as a failed + research step it reports, rather than titling clips from half a page. + +Setup details, the auth gotcha and a verified run are in [docs/brightdata.md](docs/brightdata.md). + +## Where the code runs + +Generated code never runs on the host. Two layers cover the two kinds of work. + +**Harness sandbox (Daytona).** `setup.ts` registers the provider when `DAYTONA_API_KEY` is present, +and the agent's `exec` tool then runs in a remote sandbox. Verified end to end against the running +harness: + +``` +sandbox.created sandbox_id: v1:daytona:default.e17058ee-... +exec python3 -c "print(sum(int(x)**2 for x in range(1,101)))" +tool.response {"success":true,"response":{"exitCode":0,"result":"338350\n"}} +``` + +**Media sandbox (`packages/ffmpeg-sandbox`).** ffmpeg and ffprobe are not safe to point at +agent-supplied arguments on the host, so every invocation runs in a throwaway container: +`--network none`, capped memory and CPU, `--pids-limit 256`, a non-root user, a single mounted +workspace, and a hard timeout. When its server is running, `setup.ts` registers it with the +harness as the `ffmpeg-sandbox` connector and attaches it to the agent, deferred; the +`video-editing` skill routes all raw media work through it. It exposes four tools: + +| Tool | What it does | +| --- | --- | +| `list_media` | What is in the workspace | +| `probe_media` | Duration, resolution, codecs, fps | +| `run_ffmpeg` | Runs ffmpeg with an argument array | +| `run_python` | Glue work, parsing, arithmetic | + +`run_python` is annotated `destructiveHint: true`, so the harness shows the script and waits for +approval before it runs: a script with the workspace mounted read-write can delete the source media, +and the container is not a defence against that. + +The split is deliberate. The harness sandbox is for computation the agent should not estimate; the +media sandbox is for work that must reach the media files. ## Getting started -You need [Bun](https://bun.sh) and Node 22.14+ (for the harness). +You need [Bun](https://bun.sh) and Node 22.14+ (for the harness). Docker is needed only for the +media sandbox. ```bash +git clone https://github.com/deonmenezes/edit-ai.git +cd edit-ai bun install -# 1. the harness +# 1. the harness (separate terminal) npx @truefoundry/trueforge@latest # http://localhost:8790 # 2. the timeline tools cd apps/agent && bun run start # http://localhost:8941 -# 3. wire them together (any one key is enough) -ANTHROPIC_API_KEY=sk-... bun run setup +# 3. the media sandbox (optional, needs Docker; separate terminal) +cd packages/ffmpeg-sandbox +bun run build:image && bun run start # http://localhost:8931/mcp + +# 4. wire them together (any one model key is enough) +cd apps/agent && ANTHROPIC_API_KEY=sk-... bun run setup -# 4. the editor +# 5. the editor cd ../.. && bun run dev:web # http://localhost:5173 ``` -Then ask for an edit: "Remove the silences", "Caption every video clip". +Then bring in footage. Drop any video or audio file onto the editor, or use **Import** in the Media +panel: the browser measures it with WebCodecs, uploads the bytes to the agent, and analyzes the audio +so silences, the waveform and tempo come from your file. No footage handy? + +```bash +cd apps/agent && bun scripts/make-samples.ts # writes data/samples with ffmpeg +``` + +Then ask for an edit: "Remove the silences", "Caption every video clip", "Export it at 1080p". The +export lands in `apps/agent/data/exports` as a real mp4, and the Export button turns into a download +link once it does. Without the harness running, the editor still loads with a sample timeline; the assistant panel -says it is offline. +says it is offline. That sample timeline names media it has no bytes for, so the preview says so +and export refuses until real files are imported. `setup` is idempotent, so rerun it after changing `agent.json` or adding a key. -## Stack +### Environment -- [TanStack Start](https://tanstack.com/start) + React 19, Vite, Tailwind CSS v4, shadcn/ui -- [TrueForge](https://trueforge.dev) agent harness, [MCP](https://modelcontextprotocol.io) tools -- Cloudflare Workers (via Wrangler), Bun + Turborepo monorepo +| Variable | Effect | +| --- | --- | +| `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GEMINI_API_KEY` | Registers that provider with every model in the TrueForge catalog. | +| `NVIDIA_API_KEY` | Registers NVIDIA NIM as a custom OpenAI-compatible provider. | +| `OPENAI_COMPATIBLE_BASE_URL` + `_API_KEY` + `_MODELS` (+ `_NAME`) | Registers any other OpenAI-compatible endpoint: vLLM, Ollama, a gateway. | +| `DAYTONA_API_KEY` | Configures the harness sandbox and enables it on the agent. | +| `FFMPEG_SANDBOX_URL` | Where the media sandbox serves MCP. Defaults to `http://localhost:8931/mcp`; setup attaches it only when its `/health` answers. | +| `EDITAI_MODEL` | Pins the agent's model instead of picking the best configured one. | +| `EDITAI_CONNECTORS` | Extra MCP servers to attach, comma separated. Defaults to `exa`. | +| `_MCP_HEADER` | Credential for a header-auth connector, e.g. `BRIGHT_DATA_MCP_HEADER="Authorization: Bearer ..."`. | +| `TRUEFORGE_BASE_URL` | Defaults to `http://localhost:8790`. | +| `EDITAI_AGENT_PORT` | Defaults to `8941`. | + +At least one model key is required. With none set, `setup` stops and tells you. `setup` refuses to +POST keys over plaintext HTTP to anything but localhost. + +Copy [`apps/agent/.env.example`](apps/agent/.env.example) to `apps/agent/.env` to start from a +documented set. ## Layout ``` apps/ - web/ the editor: timeline, preview, assistant panel - agent/ MCP server exposing the timeline, plus the agent definition + web/ the editor: timeline, preview, assistant panel + src/engine/ decode, composite, encode: the renderer, ported from OpenCut + media.ts probing and upload over range-requested HTTP + video-cache.ts the frame cache: forward iterator, prefetch, seek fallback + audio.ts decode, timeline mixdown, silence/peak/tempo analysis + compositor.ts one frame of the timeline, drawn in Canvas2D + exporter.ts mediabunny mux: CanvasSource + AudioBufferSource + agent/ MCP server exposing the timeline, plus the agent definition + src/tools.ts the 19 tools + src/project.ts the timeline model: split, trim, ripple delete, captions, renders + src/media.ts media names, mime types, streamed uploads + scripts/setup.ts registers models, connectors, sandbox and the agent + scripts/make-samples.ts generates real sample footage with ffmpeg +packages/ + ffmpeg-sandbox/ containerised ffmpeg/ffprobe/python MCP server +skills/ + video-editing/ ffmpeg recipes and rules the agent loads on demand +docs/ + brightdata.md connector setup and verification +site/ the landing page (static, deployed to Vercel) +tests/ the Qodo verdict parser's fixtures and tests +.github/workflows/ CI, and the Qodo merge gate ``` -See [apps/agent/README.md](apps/agent/README.md) for the tool reference and timeline semantics. - ## Scripts | Command | What it does | @@ -118,73 +371,67 @@ See [apps/agent/README.md](apps/agent/README.md) for the tool reference and time | `bun run dev` | Run every app in dev mode | | `bun run dev:web` | Run only the web app | | `bun run build` | Build every app | +| `bun run test` | Run the Bun test suites | | `bun run deploy` | Build and deploy to Cloudflare | -## Sandboxed execution +## Tests -Generated code never runs on the host. Two layers cover the two kinds of work. +**71 tests**, all passing. -**Harness sandbox (Daytona).** `setup.ts` registers the provider when -`DAYTONA_API_KEY` is present, and the agent's `exec` tool then runs in a remote -sandbox. Verified end to end against the running harness: +| Suite | Count | Covers | +| --- | --- | --- | +| `apps/agent/test/project.test.ts` | 43 | Split and trim invariants, ripple-delete arithmetic across tracks, silence removal, transcript windowing, caption merging, undo, media registration and clip placement bounds, and the render lifecycle: one claim per job, lease expiry so an abandoned render is retryable, the queued-timeline snapshot, a finished render that a late report cannot reopen, HTTP suffix-range parsing, and the path-escape guards on media names and render targets. | +| `apps/web/src/engine/audio.test.ts` | 10 | Silence detection against injected gaps and sub-threshold room tone, the peak envelope, and tempo recovered from a synthetic click track. | +| `apps/web/src/engine/compositor.test.ts` | 8 | Which clips are live at a time, source-time mapping through `sourceOffset`, and which clips reach the audio mix. | +| `apps/web/.../transcript.test.ts` | 6 | Transcript windowing and caption rendering in the editor. | +| `tests/test_qodo_verdict.py` | 8 | The merge gate's verdict parser, against the real Qodo comment bodies that broke it. | +```bash +bun run test # the Bun suites (agent + web) +cd apps/agent && bun test # the timeline model +python3 -m pytest tests/ # the Qodo verdict parser ``` -sandbox.created sandbox_id: v1:daytona:default.e17058ee-... -exec python3 -c "print(sum(int(x)**2 for x in range(1,101)))" -tool.response {"success":true,"response":{"exitCode":0,"result":"338350\n"}} -``` -**Media sandbox (`packages/ffmpeg-sandbox`).** ffmpeg and ffprobe are not safe to -point at agent-supplied arguments on the host, so every invocation runs in a -throwaway container: `--network none`, capped memory and CPU, `--pids-limit 256`, -a non-root user, a single mounted workspace, and a hard timeout. `run_python` -there is annotated `destructiveHint: true`, so the harness shows the script and -waits for approval before it runs, because a script with the workspace mounted -read-write can delete the source media and the container is not a defence -against that. - -The split is deliberate: the harness sandbox is for computation the agent should -not estimate, and the media sandbox is for work that must reach the media files. - -## Qodo Code Review Evidence - -### Merging on the review - -Qodo posts its verdict as an issue comment. It publishes no check run, no commit -status and no approving review, so GitHub's own auto-merge has nothing to gate -on. `.github/workflows/qodo-automerge.yml` is that missing gate. - -It listens for edited comments as well as created ones, because Qodo posts a -placeholder and then edits the verdict into that same comment: a gate watching -only for new comments never sees a verdict at all. On an edited event -`comment.user` is still the bot even when a person did the editing, so the -sender is checked too. - -It fails closed in every direction, because the first draft did not and Qodo -said so. It reads only the structured counter chips, never the prose: Qodo -quotes findings and diff hunks verbatim, so the words "no issues found" appear -inside reviews that are *not* clean, and any substring test on the comment body -is forgeable by the pull request's own content. It binds the verdict to the -commit Qodo footers in the comment and refuses to merge when that is no longer -the head, since a review applies to one revision and `issue_comment` runs give a -job no link to the pull request head. It treats a failure to read check state as -an error rather than as an absence of failures. And it merges with -`--match-head-commit`, so a push racing the merge is rejected by GitHub instead -of slipping in. +## Code review evidence (Qodo) + +Every pull request here is reviewed by [Qodo](https://qodo.ai). **Eight findings across three +reviews: seven were real and are fixed, one was checked and rejected with a proof.** + +### The merge gate + +Qodo posts its verdict as an issue comment. It publishes no check run, no commit status and no +approving review, so GitHub's own auto-merge has nothing to gate on. +[`.github/workflows/qodo-automerge.yml`](.github/workflows/qodo-automerge.yml) is that missing gate. + +It listens for edited comments as well as created ones, because Qodo posts a placeholder and then +edits the verdict into that same comment: a gate watching only for new comments never sees a verdict +at all. On an edited event `comment.user` is still the bot even when a person did the editing, so +the sender is checked too. + +It fails closed in every direction, because the first draft did not and Qodo said so: + +- It reads only the **structured counter chips**, never the prose. Qodo quotes findings and diff + hunks verbatim, so the words "no issues found" appear inside reviews that are *not* clean, and any + substring test on the comment body is forgeable by the pull request's own content. +- It **binds the verdict to the commit** Qodo footers in the comment and refuses to merge when that + is no longer the head, since a review applies to one revision and `issue_comment` runs give a job + no link to the pull request head. +- It treats a **failure to read check state as an error**, not as an absence of failures. +- It merges with **`--match-head-commit`**, so a push racing the merge is rejected by GitHub instead + of slipping in. ### PR #3, first review -[PR #3](https://github.com/deonmenezes/edit-ai/pull/3) was reviewed with Qodo Merge, which raised -three issues. All three were real and all three are fixed in the PR: +[PR #3](https://github.com/deonmenezes/edit-ai/pull/3) raised three issues. All three were real and +all three are fixed: 1. **Duplicate clip ids after repeated ripple deletes** (`apps/agent/src/project.ts`). The - right-hand half of a split clip took the id `${c.id}r`, so a clip cut more than once produced - the same id twice. Reproduced on the sample project: `removeSilences` left three clips sharing - `c6r` and three sharing `c7r`, which breaks clip lookup, deletion and React keys. Ids are now - allocated from the set of ids in use. Covered by three regression tests. + right-hand half of a split clip took the id `${c.id}r`, so a clip cut more than once produced the + same id twice. Reproduced on the sample project: `removeSilences` left three clips sharing `c6r` + and three sharing `c7r`, which breaks clip lookup, deletion and React keys. Ids are now allocated + from the set of ids in use. Covered by three regression tests. 2. **Export announced before the file existed.** `exportProject` committed the record, which - notifies SSE subscribers synchronously, and only then wrote the file. The write now happens - first. + notifies SSE subscribers synchronously, and only then wrote the file. The write now happens first. 3. **The session-restore effect leaked its stream on unmount**, calling `setState` on a gone component and leaving the connection open. Its cleanup now aborts the controller. @@ -201,10 +448,10 @@ rejected: - **Real:** `setup.ts` POSTs API keys to the harness, so it now refuses to do that over plaintext HTTP to anything but localhost. - **False positive:** the reviewer called the source-media bound in `trimClip` - (`end - (c.start - c.sourceOffset)`) wrong and predicted a clip could be extended to 28s instead - of 18s. The expression expands to `c.sourceOffset + (end - c.start)`, which is the correct source - time, and the code accepts exactly up to the limit and rejects one frame past it. Two tests now - pin that boundary so the correct form is not "fixed" into a broken one later. + (`end - (c.start - c.sourceOffset)`) wrong and predicted a clip could be extended to 28s instead of + 18s. The expression expands to `c.sourceOffset + (end - c.start)`, which is the correct source + time, and the code accepts exactly up to the limit and rejects one frame past it. Two tests now pin + that boundary so the correct form is not "fixed" into a broken one later. ### PR #2 review @@ -216,14 +463,43 @@ rejected: workspace is exactly what it is meant to reach. The tool is now registered with `destructiveHint: true`, so the harness stops and shows the script for approval first, the same gate the timeline's `delete_clip` and `ripple_delete` use. -- **The path guard rejected valid code.** Scanning a Python source string for `..` or a URL - rejected correct programs without adding a boundary. Path validation now applies only to - path-like arguments; for code payloads the container is the boundary. +- **The path guard rejected valid code.** Scanning a Python source string for `..` or a URL rejected + correct programs without adding a boundary. Path validation now applies only to path-like + arguments; for code payloads the container is the boundary. +## Stack + +- [TanStack Start](https://tanstack.com/start) + React 19, Vite, Tailwind CSS v4, shadcn/ui +- [TrueForge](https://trueforge.dev) agent harness, [MCP](https://modelcontextprotocol.io) tools +- [mediabunny](https://mediabunny.dev) over WebCodecs for decode, mux and encode, with the frame + cache and export pipeline ported from [OpenCut](https://github.com/opencut-app/opencut-classic) +- Cloudflare Workers (via Wrangler), Bun + Turborepo monorepo +- Landing page: hand-written static HTML in `site/`, deployed to Vercel + +## Known limits + +Worth stating plainly, because the demo does not make them obvious: + +- **Transcripts are still fixtures.** Silences, the waveform and tempo are now measured from the + decoded audio, but `transcribe_clip` reads transcript segments stored on the media rather than + running ASR. Wiring a real recognizer in is the next gap to close. +- **Rendering needs the editor open.** The agent queues a render and the browser performs it, so + `export_project` from a headless session waits for a page to claim the job. A claim carries a + 60-second lease, so a tab that closes mid-render releases the job rather than stranding it, but + something still has to pick it up. A server-side ffmpeg worker consuming the same queue would fix + it without changing any tool signature. +- **Compositing is Canvas2D.** Video, text and audio composite correctly; effects, transitions, + masks and blend modes have nowhere to live. Those need OpenCut's wgpu compositor, not this one. +- **Codecs are the browser's.** Import refuses anything Chrome cannot decode, and mp4 audio falls + back to Opus where AAC encoding is unavailable. +- Motion graphics have no dedicated timeline tool yet. Today they go through the ffmpeg sandbox or + an attached MCP server. ## Contributing -Issues and pull requests are welcome. See [CONTRIBUTING.md](.github/CONTRIBUTING.md) for setup and guidelines, and open an issue first for anything larger than a bug fix. +Issues and pull requests are welcome. See [CONTRIBUTING.md](.github/CONTRIBUTING.md) for setup and +guidelines, and open an issue first for anything larger than a bug fix. Every PR is reviewed by Qodo +before it can merge. ## License diff --git a/apps/agent/.env.example b/apps/agent/.env.example index 372fdd5..4a7c6c1 100644 --- a/apps/agent/.env.example +++ b/apps/agent/.env.example @@ -27,6 +27,13 @@ OPENAI_COMPATIBLE_MODELS= # instructions that tell it to use a sandbox become dead text. DAYTONA_API_KEY= +# Where the ffmpeg media workbench (packages/ffmpeg-sandbox) serves MCP. Setup +# probes its /health endpoint and attaches it only when it answers, so a stopped +# server is skipped with a hint rather than registered broken. Same trap as +# EDITAI_MCP_URL: setup falls back to the default only when the variable is +# absent, so leave it commented unless the server is somewhere else. +# FFMPEG_SANDBOX_URL=http://localhost:8931/mcp + # --- catalog connectors ------------------------------------------------------- # Comma separated TrueForge catalog servers to attach alongside the timeline diff --git a/apps/agent/.gitignore b/apps/agent/.gitignore index 236730b..1169630 100644 --- a/apps/agent/.gitignore +++ b/apps/agent/.gitignore @@ -1,2 +1,6 @@ +# Runtime state: the working project, imported media, renders and generated samples. +# A clone should start from the seed, not from somebody's timeline. data/project.json data/exports/ +data/media/ +data/samples/ diff --git a/apps/agent/README.md b/apps/agent/README.md index 66e8025..5e2c9a8 100644 --- a/apps/agent/README.md +++ b/apps/agent/README.md @@ -4,7 +4,7 @@ The EditAI timeline, exposed to an agent harness. This package is two things in one small server: -- **An MCP server** (`POST /mcp`) with 16 tools that read and edit a real video timeline. +- **An MCP server** (`POST /mcp`) with 19 tools that read and edit a real video timeline. - **A live project store** (`GET /project`, `GET /events`) that the web app subscribes to, so edits the agent makes appear in the editor as they happen. @@ -22,10 +22,14 @@ npx @truefoundry/trueforge@latest # http://localhost:8790 bun install bun run start # http://localhost:8941 -# 3. register models, the MCP server, the sandbox and the agent +# 3. the ffmpeg media workbench (optional, needs Docker; separate terminal) +cd ../../packages/ffmpeg-sandbox +bun run build:image && bun run start # http://localhost:8931/mcp + +# 4. register models, the MCP servers, the sandbox and the agent ANTHROPIC_API_KEY=sk-... bun run setup -# 4. the editor +# 5. the editor cd ../web && bun run dev # http://localhost:5173 ``` @@ -42,6 +46,7 @@ cd ../web && bun run dev # http://localhost:5173 | `EDITAI_MODEL` | Pins the agent's model instead of picking the best configured one. | | `EDITAI_CONNECTORS` | Extra MCP servers to attach, comma separated. Defaults to `exa` (keyless web search). | | `_MCP_HEADER` | Credential for a header-auth connector, e.g. `GITHUB_MCP_HEADER="Authorization: Bearer ghp_..."`. | +| `FFMPEG_SANDBOX_URL` | Where the ffmpeg workbench serves MCP. Defaults to `http://localhost:8931/mcp`; setup attaches it only when its `/health` answers. | | `TRUEFORGE_BASE_URL` | Defaults to `http://localhost:8790`. | | `EDITAI_AGENT_PORT` | Defaults to `8941`. | @@ -55,6 +60,7 @@ becomes a correction rather than a failed turn. | Tool | Kind | What it does | | --- | --- | --- | | `get_project` | read | The whole timeline: tracks, clips, media metadata, exports. | +| `list_media` | read | Every imported media file: duration, resolution, whether its bytes are on disk. | | `list_changes` | read | Recent edits. | | `transcribe_clip` | read | Speech inside one clip, as timed segments. | | `find_silences` | read | Silent ranges on the voice track. | @@ -63,13 +69,15 @@ becomes a correction rather than a failed turn. | `trim_clip` | write | New in/out points. | | `move_clip` | write | New start time or track. | | `set_volume` | write | Clip volume. | +| `add_clip` | write | Place an imported media file on a track. | | `add_text` | write | A title or caption. | | `add_captions` | write | Timed captions on the caption track. | | `undo` | write | Revert the last change. | | `delete_clip` | **destructive** | Remove a clip. | | `ripple_delete` | **destructive** | Remove a range from every track and close the gap. | | `remove_silences` | **destructive** | Ripple-delete every silence over a threshold. | -| `export_project` | approval | Render the timeline to a file. | +| `export_project` | approval | Queue a real render; the editor encodes it with WebCodecs. | +| `get_export` | read | Poll a queued render: progress, then the file and its byte size. | Destructive tools are published with MCP's `destructiveHint` annotation. The agent's `require_approval_for_tools: ["@destructive", "export_project"]` turns that annotation into a @@ -78,8 +86,15 @@ continues once a person allows or denies it. ## Connecting other tools -`setup.ts` attaches any server from the TrueForge catalog by name, and handles all three auth -styles: +Beyond the catalog, setup wires in this repo's own media workbench: when +`packages/ffmpeg-sandbox` is running (default `http://localhost:8931/mcp`, override with +`FFMPEG_SANDBOX_URL`), it is registered as the `ffmpeg-sandbox` connector and attached to the +agent with all tools deferred and `@destructive` approval on `run_python`. The `video-editing` +skill routes all raw media work through it. When the server is down, setup skips it with a hint +instead of registering a connector that would fail every call. + +`setup.ts` also attaches any server from the TrueForge catalog by name, and handles all three +auth styles: ```bash EDITAI_CONNECTORS=exa bun run setup # keyless, the default @@ -112,15 +127,16 @@ agent actually reaches for them. bun test ``` -12 tests over the timeline model: split/trim invariants, ripple-delete arithmetic across tracks, -silence removal, transcript windowing, caption merging, and undo. +43 tests over the timeline model: split/trim invariants, ripple-delete arithmetic across tracks, +silence removal, transcript windowing, caption merging, export lifecycle, and undo. ## Notes - The store persists to `data/project.json`, which is gitignored. `POST /project/reset` restores the sample timeline. -- `export_project` writes a JSON description of the render rather than encoding video. Wiring it to - ffmpeg is the obvious next step and does not change the agent-facing contract. +- `export_project` queues a real render. The editor claims the job, encodes it with WebCodecs, + streams the chunks back, and the file lands in `data/exports/`; the agent polls `get_export` + for progress and the finished byte size. - The MCP SDK infers handler argument types from zod shapes. That inference exhausts the TypeScript compiler on a schema set this size, so `tools.ts` registers through a narrow facade and annotates each handler's arguments explicitly. See the comment at the top of that file. diff --git a/apps/agent/data/project.json b/apps/agent/data/project.json deleted file mode 100644 index c0b503e..0000000 --- a/apps/agent/data/project.json +++ /dev/null @@ -1,337 +0,0 @@ -{ - "name": "Untitled project", - "fps": 30, - "duration": 21.1, - "tracks": [ - { - "id": "v1", - "label": "V1", - "kind": "video" - }, - { - "id": "t1", - "label": "T1", - "kind": "text" - }, - { - "id": "a1", - "label": "A1", - "kind": "audio" - }, - { - "id": "a2", - "label": "A2", - "kind": "audio" - }, - { - "id": "t2", - "label": "T2", - "kind": "text" - } - ], - "clips": [ - { - "id": "c1", - "name": "intro.mp4", - "kind": "video", - "trackId": "v1", - "start": 0, - "duration": 3.2, - "sourceOffset": 0 - }, - { - "id": "c1r", - "name": "intro.mp4", - "kind": "video", - "trackId": "v1", - "start": 3.2, - "duration": 0.9, - "sourceOffset": 4.1 - }, - { - "id": "c2", - "name": "b-roll.mp4", - "kind": "video", - "trackId": "v1", - "start": 4.1, - "duration": 4.6, - "sourceOffset": 0 - }, - { - "id": "c2r", - "name": "b-roll.mp4", - "kind": "video", - "trackId": "v1", - "start": 8.7, - "duration": 0.6, - "sourceOffset": 5.4 - }, - { - "id": "c3", - "name": "talking-head.mp4", - "kind": "video", - "trackId": "v1", - "start": 9.3, - "duration": 5.8, - "sourceOffset": 0 - }, - { - "id": "c3r", - "name": "talking-head.mp4", - "kind": "video", - "trackId": "v1", - "start": 15.1, - "duration": 6, - "sourceOffset": 7 - }, - { - "id": "c4", - "name": "Hook line", - "kind": "text", - "trackId": "t1", - "start": 0.5, - "duration": 2.7, - "sourceOffset": 0 - }, - { - "id": "c5", - "name": "Subscribe", - "kind": "text", - "trackId": "t1", - "start": 17.1, - "duration": 4, - "sourceOffset": 0 - }, - { - "id": "c6", - "name": "voiceover.wav", - "kind": "audio", - "trackId": "a1", - "start": 0, - "duration": 3.2, - "sourceOffset": 0, - "volume": 100 - }, - { - "id": "c6r", - "name": "voiceover.wav", - "kind": "audio", - "trackId": "a1", - "start": 3.2, - "duration": 5.5, - "sourceOffset": 4.1, - "volume": 100 - }, - { - "id": "c6r", - "name": "voiceover.wav", - "kind": "audio", - "trackId": "a1", - "start": 8.7, - "duration": 6.4, - "sourceOffset": 10.4, - "volume": 100 - }, - { - "id": "c6r", - "name": "voiceover.wav", - "kind": "audio", - "trackId": "a1", - "start": 15.1, - "duration": 6, - "sourceOffset": 18, - "volume": 100 - }, - { - "id": "c7", - "name": "music.mp3", - "kind": "audio", - "trackId": "a2", - "start": 0, - "duration": 3.2, - "sourceOffset": 0, - "volume": 35 - }, - { - "id": "c7r", - "name": "music.mp3", - "kind": "audio", - "trackId": "a2", - "start": 3.2, - "duration": 5.5, - "sourceOffset": 4.1, - "volume": 35 - }, - { - "id": "c7r", - "name": "music.mp3", - "kind": "audio", - "trackId": "a2", - "start": 8.7, - "duration": 6.4, - "sourceOffset": 10.4, - "volume": 35 - }, - { - "id": "c7r", - "name": "music.mp3", - "kind": "audio", - "trackId": "a2", - "start": 15.1, - "duration": 6, - "sourceOffset": 18, - "volume": 35 - }, - { - "id": "c8", - "name": "Most editors waste hours on cuts a machine should make.", - "kind": "text", - "trackId": "t2", - "start": 0.4, - "duration": 2.7, - "sourceOffset": 0 - }, - { - "id": "c9", - "name": "EditAI reads your timeline and does the boring parts.", - "kind": "text", - "trackId": "t2", - "start": 3.3, - "duration": 0.8, - "sourceOffset": 0 - }, - { - "id": "c10", - "name": "EditAI reads your timeline and does the boring parts.", - "kind": "text", - "trackId": "t2", - "start": 4.1, - "duration": 2.5, - "sourceOffset": 0 - }, - { - "id": "c11", - "name": "Silences, captions, pacing.", - "kind": "text", - "trackId": "t2", - "start": 6.7, - "duration": 1.9, - "sourceOffset": 0 - }, - { - "id": "c12", - "name": "You describe the change, it edits, you approve.", - "kind": "text", - "trackId": "t2", - "start": 8.8, - "duration": 0.5, - "sourceOffset": 0 - }, - { - "id": "c13", - "name": "You describe the change, it edits, you approve.", - "kind": "text", - "trackId": "t2", - "start": 9.3, - "duration": 3, - "sourceOffset": 0 - }, - { - "id": "c14", - "name": "Every destructive step waits for a human.", - "kind": "text", - "trackId": "t2", - "start": 12.4, - "duration": 2.6, - "sourceOffset": 0 - }, - { - "id": "c15", - "name": "It runs on any model and any tools you plug in.", - "kind": "text", - "trackId": "t2", - "start": 15.2, - "duration": 2.9, - "sourceOffset": 0 - }, - { - "id": "c16", - "name": "Subscribe if you want to see where this goes.", - "kind": "text", - "trackId": "t2", - "start": 18.3, - "duration": 2.4, - "sourceOffset": 0 - } - ], - "media": { - "intro.mp4": { - "duration": 5 - }, - "b-roll.mp4": { - "duration": 6 - }, - "talking-head.mp4": { - "duration": 13 - }, - "voiceover.wav": { - "duration": 24, - "silences": [ - { - "start": 3.2, - "end": 4.1 - }, - { - "start": 9.6, - "end": 10.4 - }, - { - "start": 16.8, - "end": 18 - } - ], - "transcript": [ - { - "start": 0.4, - "end": 3.1, - "text": "Most editors waste hours on cuts a machine should make." - }, - { - "start": 4.2, - "end": 7.5, - "text": "EditAI reads your timeline and does the boring parts." - }, - { - "start": 7.6, - "end": 9.5, - "text": "Silences, captions, pacing." - }, - { - "start": 10.5, - "end": 14, - "text": "You describe the change, it edits, you approve." - }, - { - "start": 14.1, - "end": 16.7, - "text": "Every destructive step waits for a human." - }, - { - "start": 18.1, - "end": 21, - "text": "It runs on any model and any tools you plug in." - }, - { - "start": 21.2, - "end": 23.6, - "text": "Subscribe if you want to see where this goes." - } - ] - }, - "music.mp3": { - "duration": 24, - "bpm": 120 - } - }, - "exports": [] -} \ No newline at end of file diff --git a/apps/agent/scripts/make-samples.ts b/apps/agent/scripts/make-samples.ts new file mode 100644 index 0000000..aa3c27b --- /dev/null +++ b/apps/agent/scripts/make-samples.ts @@ -0,0 +1,81 @@ +/** + * Generate real sample footage with ffmpeg, for trying EditAI without your own media. + * + * These are genuine encoded files, not fixtures: the editor decodes them, the analyzer + * measures them, and the exporter re-encodes them. The voiceover has real gaps at known + * times so silence detection has something true to find. + * + * bun scripts/make-samples.ts [outDir] + */ +import { spawn } from "node:child_process" +import { existsSync, mkdirSync } from "node:fs" +import { dirname, join } from "node:path" +import { fileURLToPath } from "node:url" + +const here = dirname(fileURLToPath(import.meta.url)) +const outDir = process.argv[2] ?? join(here, "..", "data", "samples") + +/** Gaps the voiceover really contains, so `find_silences` can be checked against the truth. */ +export const SILENCES = [ + { start: 3.2, end: 4.1 }, + { start: 9.6, end: 10.4 }, + { start: 16.8, end: 18.0 }, +] + +const gate = SILENCES.map((s) => `between(t,${s.start},${s.end})`).join("+") + +const CLIPS: { file: string; args: string[] }[] = [ + { + file: "intro.mp4", + args: ["-f", "lavfi", "-i", "testsrc2=size=1280x720:rate=30:duration=5", "-pix_fmt", "yuv420p", "-c:v", "libx264", "-preset", "veryfast"], + }, + { + file: "b-roll.mp4", + args: ["-f", "lavfi", "-i", "smptebars=size=1280x720:rate=30:duration=6", "-pix_fmt", "yuv420p", "-c:v", "libx264", "-preset", "veryfast"], + }, + { + file: "talking-head.mp4", + args: [ + "-f", "lavfi", "-i", "gradients=size=1280x720:rate=30:duration=13:c0=0x2b2440:c1=0x0e0e10", + "-pix_fmt", "yuv420p", "-c:v", "libx264", "-preset", "veryfast", + ], + }, + { + file: "voiceover.wav", + args: [ + "-f", "lavfi", + "-i", `aevalsrc='if(${gate},0,0.35*sin(2*PI*210*t)*(0.55+0.45*sin(2*PI*3.1*t)))':d=24:s=48000`, + "-c:a", "pcm_s16le", + ], + }, + { + file: "music.mp3", + // 120 BPM: a decaying click every half second, so tempo estimation has a real beat. + args: ["-f", "lavfi", "-i", "aevalsrc='0.45*sin(2*PI*760*t)*exp(-26*mod(t,0.5))':d=24:s=48000", "-c:a", "libmp3lame", "-b:a", "192k"], + }, +] + +function run(args: string[]): Promise { + return new Promise((resolve, reject) => { + const child = spawn("ffmpeg", ["-y", "-hide_banner", "-loglevel", "error", ...args], { stdio: ["ignore", "ignore", "pipe"] }) + let stderr = "" + child.stderr.on("data", (d) => (stderr += d)) + child.on("error", (err) => reject(new Error(`ffmpeg could not start: ${err.message}. Install it with: brew install ffmpeg`))) + child.on("close", (code) => (code === 0 ? resolve() : reject(new Error(stderr.trim() || `ffmpeg exited with ${code}`)))) + }) +} + +if (import.meta.main) { + mkdirSync(outDir, { recursive: true }) + for (const clip of CLIPS) { + const target = join(outDir, clip.file) + if (existsSync(target)) { + console.log(` exists ${clip.file}`) + continue + } + await run([...clip.args, target]) + console.log(` wrote ${clip.file}`) + } + console.log(`\nSample media in ${outDir}`) + console.log("Drop these onto the editor to import them. The voiceover has real silences at 3.2s, 9.6s and 16.8s.") +} diff --git a/apps/agent/scripts/setup.ts b/apps/agent/scripts/setup.ts index a33528a..e9aa6b6 100644 --- a/apps/agent/scripts/setup.ts +++ b/apps/agent/scripts/setup.ts @@ -17,6 +17,7 @@ const MCP_URL = process.env.EDITAI_MCP_URL ?? `http://localhost:${process.env.ED const AGENT_NAME = process.env.EDITAI_AGENT_NAME ?? "editai" const SKILL_REPO = process.env.EDITAI_SKILL_REPO ?? "https://github.com/deonmenezes/edit-ai" const SKILL_REF = process.env.EDITAI_SKILL_REF ?? "main" +const FFMPEG_SANDBOX_URL = process.env.FFMPEG_SANDBOX_URL ?? "http://localhost:8931/mcp" /** Catalog connectors to attach alongside the timeline tools. Keyless by default. */ const CONNECTORS = (process.env.EDITAI_CONNECTORS ?? "exa") .split(",") @@ -164,6 +165,34 @@ async function configureConnectors(): Promise<{ name: string; auth: string }[]> return attached } +/** + * The video-editing skill routes all media work through the ffmpeg-sandbox connector, + * so setup registers it whenever its server is up. It is not a catalog entry: it is this + * repo's packages/ffmpeg-sandbox, running locally in front of Docker. Registering it while + * it is down would leave the agent with a connector that fails every call, so an + * unreachable server is skipped with a hint instead. + */ +async function configureFfmpegSandbox(): Promise { + const health = FFMPEG_SANDBOX_URL.replace(/\/mcp\/?$/, "/health") + try { + const res = await fetch(health, { signal: AbortSignal.timeout(2000) }) + if (!res.ok) return false + } catch { + return false + } + await api("PUT", "/settings/mcp-servers", { + manifest: { + type: "remote", + name: "ffmpeg-sandbox", + url: FFMPEG_SANDBOX_URL, + description: + "Dockerized media workbench: list and probe media files, run ffmpeg renders, and script " + + "glue work with python. No network access; paths are relative to its workspace.", + }, + }) + return true +} + /** * Skills live in this repo, so the harness fetches them from git rather than a local path: * a remote TrueForge cannot read this machine's disk. PUT is an upsert, POST is not, so a @@ -230,10 +259,22 @@ async function pickModel(): Promise { return preferred.find((p) => names.includes(p)) ?? names[0]! } -async function upsertAgent(model: string, sandbox: boolean, connectors: { name: string; auth: string }[]) { +async function upsertAgent(model: string, sandbox: boolean, connectors: { name: string; auth: string }[], ffmpegSandbox: boolean) { const manifest = JSON.parse(readFileSync(join(here, "..", "agent.json"), "utf8")) manifest.model.name = model manifest.config.sandbox.enabled = sandbox + // The media workbench gets all its tools, deferred: the video-editing skill names them, + // so the agent finds them when a task calls for ffmpeg. run_python is published with + // destructiveHint (it can overwrite source media), and @destructive turns that into the + // same approval pause the timeline's delete tools get. + if (ffmpegSandbox && !manifest.mcp_servers.some((m: { name: string }) => m.name === "ffmpeg-sandbox")) { + manifest.mcp_servers.push({ + name: "ffmpeg-sandbox", + enable_tools: ["@all"], + preload: false, + require_approval_for_tools: ["@destructive"], + }) + } // Extra connectors are read-only and deferred: they should not enlarge the tool context // unless the agent actually reaches for research. for (const c of connectors) { @@ -272,11 +313,19 @@ const connectors = await configureConnectors() console.log( `Extra connectors: ${connectors.length ? connectors.map((c) => `${c.name} (auth: ${c.auth})`).join(", ") : "none"}`, ) +const ffmpegSandbox = await configureFfmpegSandbox() +console.log( + `ffmpeg sandbox: ${ + ffmpegSandbox + ? `attached from ${FFMPEG_SANDBOX_URL}` + : "not running (cd packages/ffmpeg-sandbox && bun run build:image && bun run start), skipped" + }`, +) const skills = await configureSkills() console.log(`Skills: ${skills.join(", ")} (from ${SKILL_REPO}@${SKILL_REF})`) const sandbox = await configureSandbox() console.log(`Sandbox: ${sandbox ? "Daytona configured, enabled on the agent" : "not configured (set DAYTONA_API_KEY to enable code execution and skills)"}`) const model = await pickModel() -const agent = await upsertAgent(model, sandbox, connectors) +const agent = await upsertAgent(model, sandbox, connectors, ffmpegSandbox) console.log(`Agent "${AGENT_NAME}" ${agent.updated ? "updated" : "created"} (id ${agent.id}) on model ${model}`) console.log(`Open ${TF} and pick "${AGENT_NAME}" in the Agents Library, or run the web app.`) diff --git a/apps/agent/src/media.ts b/apps/agent/src/media.ts new file mode 100644 index 0000000..5067400 --- /dev/null +++ b/apps/agent/src/media.ts @@ -0,0 +1,125 @@ +import { closeSync, createWriteStream, existsSync, mkdirSync, openSync, renameSync, statSync, unlinkSync, writeSync } from "node:fs" +import { randomUUID } from "node:crypto" +import { basename, dirname, extname, join } from "node:path" +import type { IncomingMessage } from "node:http" +import { pipeline } from "node:stream/promises" + +/** 2 GiB. Enough for real footage, small enough that a runaway upload cannot fill the disk. */ +export const MAX_UPLOAD_BYTES = 2 * 1024 * 1024 * 1024 + +const MIME: Record = { + ".mp4": "video/mp4", + ".mov": "video/quicktime", + ".webm": "video/webm", + ".mkv": "video/x-matroska", + ".m4v": "video/mp4", + ".wav": "audio/wav", + ".mp3": "audio/mpeg", + ".m4a": "audio/mp4", + ".aac": "audio/aac", + ".ogg": "audio/ogg", + ".flac": "audio/flac", +} + +export const mimeFor = (name: string) => MIME[extname(name).toLowerCase()] ?? "application/octet-stream" + +/** + * Media names come from the browser and end up as paths, so they are reduced to a bare + * file name. `basename` alone is not enough: a name is also used as a project key, and + * "..", an empty string or a leading dot would each produce a path that is not a file + * inside the media dir. + */ +export function safeMediaName(raw: string): string { + const name = basename(String(raw ?? "").trim()).replace(/[/\\]/g, "") + if (!name || name === "." || name === "..") throw new Error("Media name is not a usable file name.") + if (name.startsWith(".")) throw new Error("Media name cannot start with a dot.") + if (name.length > 200) throw new Error("Media name is too long.") + return name +} + +export type ByteRange = { start: number; end: number } + +/** + * Parse one HTTP byte range against a known size. + * + * The suffix form matters: `bytes=-500` means the *last* 500 bytes, not the first 500. Reading + * it as a start offset hands a decoder the head of the file when it asked for the tail, which + * is exactly where an mp4 keeps the index in a non-faststart file. + */ +export function parseRange(header: string | undefined, size: number): ByteRange | "unsatisfiable" | null { + const match = /^bytes=(\d*)-(\d*)$/.exec((header ?? "").trim()) + if (!match) return null + const [, rawStart, rawEnd] = match + if (!rawStart && !rawEnd) return "unsatisfiable" + + if (!rawStart) { + const suffix = Number(rawEnd) + if (suffix === 0) return "unsatisfiable" + return { start: Math.max(0, size - suffix), end: size - 1 } + } + + const start = Number(rawStart) + const end = rawEnd ? Math.min(Number(rawEnd), size - 1) : size - 1 + if (start >= size || end < start) return "unsatisfiable" + return { start, end } +} + +/** Stream a request body to disk, refusing anything over the cap. */ +export async function saveUpload(req: IncomingMessage, dir: string, name: string): Promise { + mkdirSync(dir, { recursive: true }) + const target = join(dir, name) + // Unique: two uploads of the same name must not write through one another's temp file. + const tmp = `${target}.${randomUUID()}.part` + let bytes = 0 + let tooBig = false + req.on("data", (chunk: Buffer) => { + bytes += chunk.length + if (bytes > MAX_UPLOAD_BYTES && !tooBig) { + tooBig = true + req.destroy(new Error(`Upload exceeds ${MAX_UPLOAD_BYTES} bytes.`)) + } + }) + try { + await pipeline(req, createWriteStream(tmp)) + } catch (err) { + if (existsSync(tmp)) unlinkSync(tmp) + throw err + } + if (bytes === 0) { + unlinkSync(tmp) + throw new Error("Upload was empty.") + } + // Rename last so a reader never sees a half-written file under the real name. + renameSync(tmp, target) + return statSync(target).size +} + +/** Collect a bounded request body. Used for render chunks, which are capped by the writer. */ +export function readBody(req: IncomingMessage, maxBytes: number): Promise { + return new Promise((resolve, reject) => { + const parts: Buffer[] = [] + let size = 0 + req.on("data", (chunk: Buffer) => { + size += chunk.length + if (size > maxBytes) { + req.destroy() + reject(new Error(`Chunk exceeds ${maxBytes} bytes.`)) + return + } + parts.push(chunk) + }) + req.on("end", () => resolve(Buffer.concat(parts))) + req.on("error", reject) + }) +} + +/** Write a chunk at its byte offset, creating the file if this is the first one. */ +export function writeChunkAt(path: string, position: number, data: Buffer): void { + mkdirSync(dirname(path), { recursive: true }) + const fd = openSync(path, existsSync(path) ? "r+" : "w+") + try { + writeSync(fd, data, 0, data.length, position) + } finally { + closeSync(fd) + } +} diff --git a/apps/agent/src/project.ts b/apps/agent/src/project.ts index 38c650e..e33da9b 100644 --- a/apps/agent/src/project.ts +++ b/apps/agent/src/project.ts @@ -1,5 +1,5 @@ -import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs" -import { dirname } from "node:path" +import { existsSync, mkdirSync, readFileSync, renameSync, statSync, unlinkSync, writeFileSync } from "node:fs" +import { dirname, join } from "node:path" export type ClipKind = "video" | "text" | "audio" @@ -31,15 +31,41 @@ export type MediaInfo = { transcript?: Segment[] /** beats per minute, for music */ bpm?: number + /** Set once real bytes are on disk. File name inside the media dir. */ + file?: string + width?: number + height?: number + fps?: number + hasAudio?: boolean + sizeBytes?: number + /** Peak envelope over the whole file, 0..1, for the timeline waveform. */ + peaks?: number[] + /** Where silences/peaks came from: absent means they were never measured. */ + analyzedAt?: string } +export type ExportStatus = "pending" | "rendering" | "done" | "failed" + export type ExportRecord = { id: string format: string resolution: string + width: number + height: number + fps: number createdAt: string durationSeconds: number file: string + status: ExportStatus + /** 0..1 while rendering. */ + progress?: number + /** Real bytes on disk, only once status is "done". */ + sizeBytes?: number + completedAt?: string + error?: string + /** When a worker took the job, and when it last showed a sign of life. */ + claimedAt?: string + heartbeatAt?: string } export type Project = { @@ -112,6 +138,11 @@ export function seedProject(): Project { } } +/** A project with the standard track layout and nothing on it, for starting from real footage. */ +export function emptyProject(): Project { + return { ...seedProject(), clips: [], media: {}, duration: 0, exports: [] } +} + export class ProjectStore { private project: Project private snapshots: Project[] = [] @@ -119,10 +150,21 @@ export class ProjectStore { private listeners = new Set<(p: Project, change: Change | null) => void>() revision = 0 - constructor(private file?: string) { + constructor( + private file?: string, + /** Where media bytes live. Given, the store can tell a registered file from a present one. */ + private mediaDir?: string, + ) { this.project = file && existsSync(file) ? (JSON.parse(readFileSync(file, "utf8")) as Project) : seedProject() } + /** Media is only usable once its bytes are actually on disk, not merely named. */ + hasBytes(name: string): boolean { + const info = this.project.media[name] + if (!info?.file) return false + return this.mediaDir ? existsSync(join(this.mediaDir, info.file)) : true + } + get(): Project { return structuredClone(this.project) } @@ -136,10 +178,10 @@ export class ProjectStore { return () => this.listeners.delete(fn) } - reset() { + reset(opts: { empty?: boolean } = {}) { this.snapshots = [] this.changes = [] - this.project = seedProject() + this.project = opts.empty ? emptyProject() : seedProject() this.persist(null) } @@ -404,22 +446,279 @@ export class ProjectStore { return { track, clips: this.project.clips.filter((c) => ids.includes(c.id)) } } - exportProject(format: string, resolution: string, dir: string): ExportRecord { + // ---- real media ----------------------------------------------------------- + + /** + * Record media whose bytes are on disk. Metadata is measured by the editor (WebCodecs) + * rather than guessed here, so the agent server needs no ffprobe of its own. + */ + registerMedia(name: string, info: Omit): MediaInfo { + if (!name.trim()) throw new Error("Media needs a name.") + this.commit("import_media", `Imported ${name} (${round(info.duration)}s)`, (p) => { + p.media[name] = { ...p.media[name], ...info } + }) + return this.project.media[name]! + } + + /** Attach measurements taken from the decoded audio: silences, peak envelope, tempo. */ + setMediaAnalysis(name: string, analysis: { silences?: Segment[]; peaks?: number[]; bpm?: number; transcript?: Segment[] }): MediaInfo { + const media = this.project.media[name] + if (!media) throw new Error(`No media named "${name}".`) + const counted = analysis.silences ? `${analysis.silences.length} silences` : "waveform" + this.commit("analyze_media", `Analyzed ${name}: ${counted}`, (p) => { + p.media[name] = { ...p.media[name]!, ...analysis, analyzedAt: new Date().toISOString() } + }) + return this.project.media[name]! + } + + /** Put imported media on the timeline. Unlike addTextClip this needs a real source file. */ + addClip(opts: { name: string; trackId: string; start: number; duration?: number; sourceOffset?: number }): Clip { + const track = this.track(opts.trackId) + if (track.kind === "text") throw new Error(`${track.label} is a text track; use add_text instead.`) + const media = this.project.media[opts.name] + if (!media) throw new Error(`No media named "${opts.name}". Call list_media to see what has been imported.`) + if (!this.hasBytes(opts.name)) throw new Error(`"${opts.name}" has no media on disk yet, so it cannot be placed on the timeline.`) + const sourceOffset = opts.sourceOffset ?? 0 + if (sourceOffset < 0 || sourceOffset >= media.duration) { + throw new Error(`sourceOffset ${sourceOffset}s is outside ${opts.name} (0s to ${media.duration}s).`) + } + const duration = opts.duration ?? media.duration - sourceOffset + if (duration <= 0) throw new Error("Duration must be positive.") + if (sourceOffset + duration > media.duration + 1e-6) { + throw new Error(`${opts.name} only has ${round(media.duration - sourceOffset)}s left after a ${sourceOffset}s offset.`) + } + if (opts.start < 0) throw new Error("Clips cannot start before 0s.") + let id = "" + this.commit("add_clip", `Added ${opts.name} at ${round(opts.start)}s on ${track.label}`, (p) => { + id = this.newId("c") + p.clips.push({ + id, + name: opts.name, + kind: track.kind, + trackId: track.id, + start: opts.start, + duration, + sourceOffset, + ...(track.kind === "audio" ? { volume: 100 } : {}), + }) + }) + return this.clip(id) + } + + // ---- export --------------------------------------------------------------- + + /** + * Queue a render. The encode happens in the editor, which owns the decoders, so this + * only creates the job; the file appears when the editor posts the bytes back. + */ + requestExport(format: string, resolution: string, dir: string): ExportRecord { + assertRenderTarget(format, resolution) + const missing = this.missingMedia() + if (missing.length > 0) { + throw new Error( + `Cannot render: ${missing.join(", ")} ${missing.length === 1 ? "has" : "have"} no media on disk. ` + + `Import real footage first, or remove those clips.`, + ) + } + const { width, height } = resolutionToSize(resolution) + const used = new Set(this.project.exports.map((e) => e.id)) + let n = this.project.exports.length + 1 + while (used.has(`exp${n}`)) n++ const rec: ExportRecord = { - id: `exp${this.project.exports.length + 1}`, + id: `exp${n}`, format, resolution, + width, + height, + fps: this.project.fps, createdAt: new Date().toISOString(), durationSeconds: this.project.duration, - file: `${dir}/${this.project.name.replace(/\s+/g, "-").toLowerCase()}-${resolution}.${format}`, + file: join(dir, `${slug(this.project.name)}-${resolution}-exp${n}.${format}`), + status: "pending", + progress: 0, } - // Written before the commit: commit notifies subscribers synchronously, and a client that - // reacts to the new export record must not find the file missing. mkdirSync(dir, { recursive: true }) - writeFileSync(rec.file, JSON.stringify({ export: rec, project: this.project }, null, 2)) - this.commit("export_project", `Exported ${resolution} ${format}`, (p) => { + // The render must be of the timeline as it was approved, not of whatever it has become by + // the time an editor picks the job up, so the project is frozen here beside the job. + writeFileSync(snapshotPath(rec), JSON.stringify(this.project, null, 2)) + this.commit("export_project", `Queued a ${resolution} ${format} render`, (p) => { p.exports.push(rec) }) return rec } + + /** The timeline as it stood when a render was queued. */ + exportSnapshot(id: string): Project { + const rec = this.getExport(id) + const path = snapshotPath(rec) + if (!existsSync(path)) throw new Error(`Export ${id} has no snapshot; it cannot be rendered faithfully.`) + return JSON.parse(readFileSync(path, "utf8")) as Project + } + + /** Clips whose source media has no bytes on disk. Nothing can be rendered from those. */ + missingMedia(): string[] { + const names = new Set() + for (const c of this.project.clips) { + if (c.kind === "text") continue + if (!this.hasBytes(c.name)) names.add(c.name) + } + return [...names] + } + + getExport(id: string): ExportRecord { + const rec = this.project.exports.find((e) => e.id === id) + if (!rec) throw new Error(`No export with id "${id}".`) + return structuredClone(rec) + } + + /** The oldest render an editor could take: never started, or abandoned mid-flight. */ + pendingExport(): ExportRecord | null { + return structuredClone(this.project.exports.find(isClaimable) ?? null) + } + + private updateExport(id: string, fn: (rec: ExportRecord) => void, op: string, summary: string) { + if (!this.project.exports.some((e) => e.id === id)) throw new Error(`No export with id "${id}".`) + this.commit(op, summary, (p) => { + fn(p.exports.find((e) => e.id === id)!) + }) + return this.getExport(id) + } + + /** + * Take a queued render. Returns null if it is already claimed, which is how two editors + * open on the same project avoid both encoding it: only one claim can win. + */ + claimExport(id: string): ExportRecord | null { + const rec = this.project.exports.find((e) => e.id === id) + if (!rec || !isClaimable(rec)) return null + const retry = rec.status === "rendering" + const now = new Date().toISOString() + return this.updateExport( + id, + (r) => { + r.status = "rendering" + r.progress = 0 + r.claimedAt = now + r.heartbeatAt = now + }, + "export_claimed", + retry ? `Reclaimed abandoned render ${id}` : `Started rendering ${id}`, + ) + } + + /** Queued, or claimed by a worker that has gone quiet for longer than the lease. */ + claimableExports(): ExportRecord[] { + return this.project.exports.filter(isClaimable).map((e) => structuredClone(e)) + } + + setExportProgress(id: string, progress: number) { + // A finished render must not be reopened by a straggling progress report. + const current = this.getExport(id) + if (current.status !== "rendering") return current + const clamped = Math.max(0, Math.min(1, progress)) + return this.updateExport( + id, + (rec) => { + rec.progress = clamped + rec.heartbeatAt = new Date().toISOString() + }, + "export_progress", + `Rendering ${id}: ${Math.round(clamped * 100)}%`, + ) + } + + /** + * Finish a render whose bytes are already on disk at `uploadedPath`. + * + * A path rather than a buffer: a 4K render is gigabytes, and reading it into the server only + * to write it out again would be the one place this design needs the whole file in memory. + */ + completeExport(id: string, uploadedPath: string) { + const rec = this.getExport(id) + if (!existsSync(uploadedPath)) throw new Error(`No uploaded file at ${uploadedPath}.`) + // Moved before the commit: commit notifies subscribers synchronously, and a client that + // reacts to the finished export must not find the file missing. + mkdirSync(dirname(rec.file), { recursive: true }) + renameSync(uploadedPath, rec.file) + const sizeBytes = statSync(rec.file).size + // The snapshot exists so an abandoned render can be retried faithfully. Done is terminal, + // so it has nothing left to serve. + const snapshot = snapshotPath(rec) + if (existsSync(snapshot)) unlinkSync(snapshot) + return this.updateExport( + id, + (r) => { + r.status = "done" + r.progress = 1 + r.sizeBytes = sizeBytes + r.completedAt = new Date().toISOString() + delete r.error + }, + "export_done", + `Rendered ${rec.resolution} ${rec.format} (${(sizeBytes / 1e6).toFixed(1)} MB)`, + ) + } + + failExport(id: string, error: string) { + // Same reason: a late failure from a losing worker must not bury a finished file. + const current = this.getExport(id) + if (current.status === "done") return current + return this.updateExport( + id, + (rec) => { + rec.status = "failed" + rec.error = error + }, + "export_failed", + `Render ${id} failed: ${error}`, + ) + } +} + +/** + * How long a claimed render may go silent before another editor may take it over. + * + * Without this a tab that is closed mid-render strands the job in `rendering` forever: only + * the worker that claimed it can report failure, and it is gone. + */ +export const RENDER_LEASE_MS = 60_000 + +export function isClaimable(rec: ExportRecord, now = Date.now()): boolean { + if (rec.status === "pending") return true + if (rec.status !== "rendering") return false + const last = Date.parse(rec.heartbeatAt ?? rec.claimedAt ?? rec.createdAt) + return Number.isFinite(last) && now - last > RENDER_LEASE_MS +} + +const snapshotPath = (rec: ExportRecord) => `${rec.file}.project.json` + +/** Render sizes are 16:9, matching the editor's preview. */ +const SIZES: Record = { + "720p": { width: 1280, height: 720 }, + "1080p": { width: 1920, height: 1080 }, + "4k": { width: 3840, height: 2160 }, +} + +const FORMATS = new Set(["mp4", "webm"]) + +export function resolutionToSize(resolution: string): { width: number; height: number } { + const size = SIZES[resolution] + if (!size) throw new Error(`Unknown resolution "${resolution}". Use one of: ${Object.keys(SIZES).join(", ")}.`) + return size +} + +/** + * Both of these end up inside the output path. The MCP tool constrains them with zod, but the + * HTTP route does not, and `join` happily normalizes `../` out of a directory, so they are + * checked against the allowed sets here where the path is actually built. + */ +function assertRenderTarget(format: string, resolution: string) { + if (!FORMATS.has(format)) throw new Error(`Unknown format "${format}". Use one of: ${[...FORMATS].join(", ")}.`) + resolutionToSize(resolution) +} + +/** The project name is part of the file name, so it is reduced to a safe slug. */ +function slug(name: string): string { + const cleaned = name.replace(/[^a-zA-Z0-9-_ ]/g, "").trim().replace(/\s+/g, "-").toLowerCase() + return cleaned.slice(0, 60) || "project" } diff --git a/apps/agent/src/server.ts b/apps/agent/src/server.ts index e2cecf2..67a4e5b 100644 --- a/apps/agent/src/server.ts +++ b/apps/agent/src/server.ts @@ -1,22 +1,35 @@ import { StreamableHTTPServerTransport } from "@modelcontextprotocol/sdk/server/streamableHttp.js" import cors from "cors" import express from "express" +import { createReadStream, existsSync, statSync, unlinkSync } from "node:fs" import { dirname, join } from "node:path" import { fileURLToPath } from "node:url" +import { mimeFor, parseRange, readBody, safeMediaName, saveUpload, writeChunkAt } from "./media.ts" import { ProjectStore } from "./project.ts" import { buildServer } from "./tools.ts" const here = dirname(fileURLToPath(import.meta.url)) const DATA_DIR = process.env.EDITAI_DATA_DIR ?? join(here, "..", "data") +const MEDIA_DIR = join(DATA_DIR, "media") +const EXPORTS_DIR = join(DATA_DIR, "exports") const PORT = Number(process.env.EDITAI_AGENT_PORT ?? 8941) -const store = new ProjectStore(join(DATA_DIR, "project.json")) +/** Where an in-flight encode accumulates before it is moved to its final name. */ +const renderPath = (id: string) => join(EXPORTS_DIR, `${id}.render`) + +const store = new ProjectStore(join(DATA_DIR, "project.json"), MEDIA_DIR) const app = express() app.use(cors({ origin: true })) -app.use(express.json({ limit: "2mb" })) + +/** Uploads stream straight to disk, so they must not be buffered by the JSON parser first. */ +const isUpload = (url: string) => /^\/(media\/[^/]+|exports\/[^/]+\/chunk)$/.test(url.split("?")[0]!) +app.use((req, res, next) => (req.method === "POST" && isUpload(req.url) ? next() : express.json({ limit: "8mb" })(req, res, next))) + +const fail = (res: express.Response, status: number, err: unknown) => + res.status(status).json({ error: err instanceof Error ? err.message : String(err) }) app.get("/", (_req, res) => { - res.type("text/plain").send("EditAI agent server. MCP at POST /mcp. Project at GET /project, live at GET /events.") + res.type("text/plain").send("EditAI agent server. MCP at POST /mcp. Project at GET /project, live at GET /events, media at GET /media.") }) app.get("/healthz", (_req, res) => res.json({ ok: true, revision: store.revision })) @@ -25,11 +38,183 @@ app.get("/project", (_req, res) => { res.json({ project: store.get(), revision: store.revision, changes: store.listChanges(20) }) }) -app.post("/project/reset", (_req, res) => { - store.reset() +app.post("/project/reset", (req, res) => { + store.reset({ empty: req.body?.empty === true }) res.json({ project: store.get(), revision: store.revision }) }) +// ---- media ---------------------------------------------------------------- + +app.get("/media", (_req, res) => res.json({ media: store.get().media, dir: MEDIA_DIR })) + +/** + * The editor measures media with WebCodecs and posts the bytes here, so the server needs + * no decoder of its own. Metadata rides in the query string; the body is the file. + */ +app.post("/media/:name", async (req, res) => { + try { + const name = safeMediaName(req.params.name) + const q = req.query as Record + const duration = Number(q.duration) + if (!Number.isFinite(duration) || duration <= 0) throw new Error("A positive ?duration= in seconds is required.") + const sizeBytes = await saveUpload(req, MEDIA_DIR, name) + const num = (v: string | undefined) => (v === undefined || v === "" ? undefined : Number(v)) + const info = store.registerMedia(name, { + duration, + file: name, + sizeBytes, + width: num(q.width), + height: num(q.height), + fps: num(q.fps), + hasAudio: q.hasAudio === undefined ? undefined : q.hasAudio === "true", + }) + res.json({ media: info, name }) + } catch (err) { + fail(res, 400, err) + } +}) + +app.post("/media/:name/analysis", (req, res) => { + try { + res.json({ media: store.setMediaAnalysis(safeMediaName(req.params.name), req.body ?? {}) }) + } catch (err) { + fail(res, 400, err) + } +}) + +app.get("/media/:name", (req, res) => { + try { + const name = safeMediaName(req.params.name) + const file = join(MEDIA_DIR, name) + if (!existsSync(file)) return fail(res, 404, new Error(`No media file named "${name}".`)) + const { size } = statSync(file) + res.setHeader("Content-Type", mimeFor(name)) + res.setHeader("Accept-Ranges", "bytes") + // Range support keeps a decoder from pulling the whole file to read one atom. + const range = parseRange(req.headers.range, size) + if (range === "unsatisfiable") { + res.setHeader("Content-Range", `bytes */${size}`) + return res.status(416).end() + } + if (range) { + res.status(206) + res.setHeader("Content-Range", `bytes ${range.start}-${range.end}/${size}`) + res.setHeader("Content-Length", String(range.end - range.start + 1)) + return createReadStream(file, { start: range.start, end: range.end }).pipe(res) + } + res.setHeader("Content-Length", String(size)) + createReadStream(file).pipe(res) + } catch (err) { + fail(res, 400, err) + } +}) + +// ---- renders -------------------------------------------------------------- + +/** The Export button queues the same job the agent's export_project tool does. */ +app.post("/exports", (req, res) => { + try { + const format = String(req.body?.format ?? "mp4") + const resolution = String(req.body?.resolution ?? "1080p") + res.json({ export: store.requestExport(format, resolution, EXPORTS_DIR) }) + } catch (err) { + fail(res, 400, err) + } +}) + +/** The editor claims the oldest queued render and encodes it. */ +app.get("/exports/pending", (_req, res) => res.json({ export: store.pendingExport() })) + +/** Claiming is what stops two open editors from encoding the same job twice. */ +app.post("/exports/:id/claim", (req, res) => { + try { + const claimed = store.claimExport(req.params.id) + if (!claimed) return res.status(409).json({ error: `Export ${req.params.id} is not claimable.`, export: store.getExport(req.params.id) }) + // A reclaim starts from nothing. Writing a fresh encode over an abandoned one leaves any + // trailing bytes of the longer attempt behind, which is a corrupt file rather than a retry. + const partial = renderPath(claimed.id) + if (existsSync(partial)) unlinkSync(partial) + res.json({ export: claimed }) + } catch (err) { + fail(res, 404, err) + } +}) + +app.post("/exports/:id/progress", (req, res) => { + try { + res.json({ export: store.setExportProgress(req.params.id, Number(req.body?.progress)) }) + } catch (err) { + fail(res, 400, err) + } +}) + +app.post("/exports/:id/failed", (req, res) => { + try { + // A half-written render is not worth keeping; the retry starts from an empty file. + const partial = renderPath(req.params.id) + if (existsSync(partial)) unlinkSync(partial) + res.json({ export: store.failExport(req.params.id, String(req.body?.error ?? "Unknown render error")) }) + } catch (err) { + fail(res, 400, err) + } +}) + +/** The timeline as it was when the render was queued. Workers render this, not the live project. */ +app.get("/exports/:id/project", (req, res) => { + try { + res.json({ project: store.exportSnapshot(req.params.id) }) + } catch (err) { + fail(res, 404, err) + } +}) + +/** + * One slice of the encode, written at its byte offset. + * + * The muxer emits chunks as it goes and revisits earlier offsets to finish the index, so this + * takes a position rather than appending. Uploading as it encodes is what keeps a long render + * off the browser's heap. + */ +const CHUNK_LIMIT = 32 * 1024 * 1024 + +app.post("/exports/:id/chunk", async (req, res) => { + try { + const rec = store.getExport(req.params.id) + const position = Number((req.query as Record).position) + if (!Number.isInteger(position) || position < 0) throw new Error("A non-negative integer ?position= is required.") + const data = await readBody(req, CHUNK_LIMIT) + if (data.length === 0) throw new Error("Chunk was empty.") + writeChunkAt(renderPath(rec.id), position, data) + res.json({ ok: true, position, bytes: data.length }) + } catch (err) { + fail(res, 400, err) + } +}) + +/** The encode is complete: move what the chunks built into place. */ +app.post("/exports/:id/finish", (req, res) => { + try { + const rec = store.getExport(req.params.id) + res.json({ export: store.completeExport(rec.id, renderPath(rec.id)) }) + } catch (err) { + fail(res, 400, err) + } +}) + +app.get("/exports/:id/file", (req, res) => { + try { + const rec = store.getExport(req.params.id) + if (rec.status !== "done" || !existsSync(rec.file)) return fail(res, 404, new Error(`Export ${rec.id} has no file yet.`)) + res.setHeader("Content-Type", mimeFor(rec.file)) + res.setHeader("Content-Disposition", `attachment; filename="${rec.file.split("/").pop()}"`) + createReadStream(rec.file).pipe(res) + } catch (err) { + fail(res, 404, err) + } +}) + +// ---- live ----------------------------------------------------------------- + app.get("/events", (req, res) => { res.setHeader("Content-Type", "text/event-stream") res.setHeader("Cache-Control", "no-cache") @@ -47,7 +232,7 @@ app.get("/events", (req, res) => { app.post("/mcp", async (req, res) => { try { - const server = buildServer(store, join(DATA_DIR, "exports")) + const server = buildServer(store, EXPORTS_DIR) const transport = new StreamableHTTPServerTransport({ sessionIdGenerator: undefined }) res.on("close", () => { transport.close() @@ -68,5 +253,6 @@ app.all("/mcp", (_req, res) => { }) app.listen(PORT, () => { - console.log(`[editai-agent] listening on http://localhost:${PORT} (MCP: /mcp, project: /project, events: /events)`) + console.log(`[editai-agent] listening on http://localhost:${PORT} (MCP: /mcp, project: /project, media: /media, events: /events)`) + console.log(`[editai-agent] media dir: ${MEDIA_DIR}`) }) diff --git a/apps/agent/src/tools.ts b/apps/agent/src/tools.ts index 511fbae..71dca4f 100644 --- a/apps/agent/src/tools.ts +++ b/apps/agent/src/tools.ts @@ -67,6 +67,27 @@ export function buildServer(store: ProjectStore, exportsDir: string) { handler(() => run(() => store.get())), ) + server.registerTool( + "list_media", + { + title: "List imported media", + description: + "Every media file the user has imported, with its real duration, resolution, frame rate and whether its bytes are on disk. Only media listed here can be placed on the timeline or rendered.", + inputSchema: {}, + annotations: READ, + }, + handler(() => + run(() => { + const project = store.get() + return Object.entries(project.media).map(([name, info]) => ({ + name, + ...info, + onDisk: store.hasBytes(name), + })) + }), + ), + ) + server.registerTool( "list_changes", { @@ -173,6 +194,38 @@ export function buildServer(store: ProjectStore, exportsDir: string) { handler(({ clip_id, volume }: { clip_id: string; volume: number }) => run(() => store.setVolume(clip_id, volume))), ) + server.registerTool( + "add_clip", + { + title: "Add a media clip", + description: + "Place imported media on a video or audio track. Defaults to the whole file. Use list_media first: only media with bytes on disk can be placed.", + inputSchema: { + name: z.string().describe("Media file name from list_media, e.g. interview.mp4"), + track_id: z.string().describe("Video or audio track id, e.g. v1 or a1"), + start: z.number().min(0).describe("Timeline seconds"), + duration: z.number().min(0.1).optional().describe("Defaults to the rest of the file"), + source_offset: z.number().min(0).optional().describe("Seconds into the source to start from"), + }, + annotations: WRITE, + }, + handler( + ({ + name, + track_id, + start, + duration, + source_offset, + }: { + name: string + track_id: string + start: number + duration?: number + source_offset?: number + }) => run(() => store.addClip({ name, trackId: track_id, start, duration, sourceOffset: source_offset })), + ), + ) + server.registerTool( "add_text", { @@ -262,17 +315,33 @@ export function buildServer(store: ProjectStore, exportsDir: string) { "export_project", { title: "Export the project", - description: "Render the timeline to a file. Always asks the user to approve first.", + description: + "Queue a real render of the timeline. The editor encodes it with WebCodecs and writes the file, so this returns immediately with a pending export; poll get_export until it is done, then report the real file and size. Fails if any clip's media is missing. Always asks the user to approve first.", inputSchema: { - format: z.enum(["mp4", "mov", "webm"]).default("mp4"), + format: z.enum(["mp4", "webm"]).default("mp4"), resolution: z.enum(["720p", "1080p", "4k"]).default("1080p"), }, annotations: WRITE, }, handler(({ format, resolution }: { format: string; resolution: string }) => - run(() => store.exportProject(format, resolution, exportsDir)), + run(() => ({ + ...store.requestExport(format, resolution, exportsDir), + note: "Rendering happens in the editor. Poll get_export with this id until status is done or failed.", + })), ), ) + server.registerTool( + "get_export", + { + title: "Check a render", + description: + "The status of a queued render: pending, rendering (with progress), done (with the file path and real byte size) or failed (with the error).", + inputSchema: { export_id: z.string().describe("Export id from export_project, e.g. exp1") }, + annotations: READ, + }, + handler(({ export_id }: { export_id: string }) => run(() => store.getExport(export_id))), + ) + return mcp } diff --git a/apps/agent/test/project.test.ts b/apps/agent/test/project.test.ts index 46b9c51..39f8f5b 100644 --- a/apps/agent/test/project.test.ts +++ b/apps/agent/test/project.test.ts @@ -1,5 +1,20 @@ import { describe, expect, test } from "bun:test" -import { ProjectStore } from "../src/project.ts" +import { isClaimable, ProjectStore, RENDER_LEASE_MS } from "../src/project.ts" +import { parseRange, safeMediaName } from "../src/media.ts" +import { randomUUID } from "node:crypto" +import { mkdirSync, writeFileSync } from "node:fs" +import { tmpdir } from "node:os" +import { join } from "node:path" + +const EXPORT_DIR = join(tmpdir(), "editai-test-exports") + +/** Stand in for a render the editor has already streamed to disk. */ +function uploaded(bytes: Uint8Array): string { + mkdirSync(EXPORT_DIR, { recursive: true }) + const path = join(EXPORT_DIR, `upload-${randomUUID()}`) + writeFileSync(path, bytes) + return path +} describe("ProjectStore", () => { test("split keeps media in sync", () => { @@ -166,3 +181,219 @@ describe("trim respects the source media after a split", () => { expect(() => s.trimClip(right.id, { end: 24.5 })).toThrow(/source media/) }) }) + +describe("real media", () => { + /** No media dir given, so registered media counts as present: file-system checks are covered separately. */ + const withMedia = () => { + const s = new ProjectStore() + s.reset({ empty: true }) + s.registerMedia("clip.mp4", { duration: 12, file: "clip.mp4", width: 1920, height: 1080, fps: 30, hasAudio: true }) + return s + } + + test("an empty project starts with tracks but nothing on them", () => { + const s = new ProjectStore() + s.reset({ empty: true }) + expect(s.get().clips).toEqual([]) + expect(s.get().media).toEqual({}) + expect(s.get().duration).toBe(0) + }) + + test("registering media records what was measured", () => { + const s = withMedia() + expect(s.get().media["clip.mp4"]).toMatchObject({ duration: 12, width: 1920, hasAudio: true }) + expect(s.hasBytes("clip.mp4")).toBe(true) + }) + + test("a clip defaults to the rest of the file", () => { + const s = withMedia() + const clip = s.addClip({ name: "clip.mp4", trackId: "v1", start: 2, sourceOffset: 4 }) + expect(clip.duration).toBe(8) + expect(clip.start).toBe(2) + expect(s.get().duration).toBe(10) + }) + + test("a clip cannot run past the end of its source", () => { + const s = withMedia() + expect(() => s.addClip({ name: "clip.mp4", trackId: "v1", start: 0, sourceOffset: 8, duration: 6 })).toThrow(/only has 4s/) + }) + + test("media that was never imported cannot be placed", () => { + const s = withMedia() + expect(() => s.addClip({ name: "ghost.mp4", trackId: "v1", start: 0 })).toThrow(/No media named/) + }) + + test("audio cannot be placed on a text track", () => { + const s = withMedia() + expect(() => s.addClip({ name: "clip.mp4", trackId: "t1", start: 0 })).toThrow(/text track/) + }) + + test("analysis attaches measurements to the media", () => { + const s = withMedia() + s.setMediaAnalysis("clip.mp4", { silences: [{ start: 1, end: 2 }], peaks: [0.1, 0.9], bpm: 120 }) + const media = s.get().media["clip.mp4"]! + expect(media.silences).toEqual([{ start: 1, end: 2 }]) + expect(media.bpm).toBe(120) + expect(media.analyzedAt).toBeTruthy() + }) +}) + +describe("export lifecycle", () => { + const ready = () => { + const s = new ProjectStore() + s.reset({ empty: true }) + s.registerMedia("clip.mp4", { duration: 10, file: "clip.mp4", width: 1920, height: 1080 }) + s.addClip({ name: "clip.mp4", trackId: "v1", start: 0 }) + return s + } + + test("a render is queued, not written", () => { + const rec = ready().requestExport("mp4", "1080p", EXPORT_DIR) + expect(rec.status).toBe("pending") + expect(rec.width).toBe(1920) + expect(rec.sizeBytes).toBeUndefined() + }) + + test("resolution decides the frame size", () => { + const s = ready() + expect(s.requestExport("mp4", "720p", EXPORT_DIR)).toMatchObject({ width: 1280, height: 720 }) + expect(s.requestExport("mp4", "4k", EXPORT_DIR)).toMatchObject({ width: 3840, height: 2160 }) + }) + + test("the sample timeline cannot be rendered until real media is imported", () => { + const s = new ProjectStore() + expect(s.missingMedia().sort()).toEqual(["b-roll.mp4", "intro.mp4", "music.mp3", "talking-head.mp4", "voiceover.wav"]) + expect(() => s.requestExport("mp4", "1080p", EXPORT_DIR)).toThrow(/no media on disk/) + }) + + test("only one worker can claim a render", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + expect(s.claimExport(rec.id)?.status).toBe("rendering") + expect(s.claimExport(rec.id)).toBeNull() + }) + + test("a finished render is not reopened by a late progress or failure report", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + s.claimExport(rec.id) + s.completeExport(rec.id, uploaded(new Uint8Array([1, 2, 3, 4]))) + expect(s.getExport(rec.id)).toMatchObject({ status: "done", sizeBytes: 4 }) + s.setExportProgress(rec.id, 0.5) + s.failExport(rec.id, "too late") + expect(s.getExport(rec.id)).toMatchObject({ status: "done", sizeBytes: 4 }) + }) + + test("a failed render keeps its error", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + s.claimExport(rec.id) + expect(s.failExport(rec.id, "codec unsupported")).toMatchObject({ status: "failed", error: "codec unsupported" }) + }) +}) + +describe("render safety", () => { + const ready = () => { + const s = new ProjectStore() + s.reset({ empty: true }) + s.registerMedia("clip.mp4", { duration: 10, file: "clip.mp4", width: 1920, height: 1080 }) + s.addClip({ name: "clip.mp4", trackId: "v1", start: 0 }) + return s + } + + test("format and resolution cannot escape the exports directory", () => { + const s = ready() + // These become path components, and join() would normalize the traversal away. + expect(() => s.requestExport("mp4", "../../../outside", EXPORT_DIR)).toThrow(/Unknown resolution/) + expect(() => s.requestExport("../../evil", "1080p", EXPORT_DIR)).toThrow(/Unknown format/) + expect(s.requestExport("mp4", "1080p", EXPORT_DIR).file.startsWith(EXPORT_DIR)).toBe(true) + }) + + test("the project name cannot escape it either", () => { + const s = ready() + s.get() // touch, then rename through the persisted model + const store = new ProjectStore() + store.reset({ empty: true }) + store.registerMedia("clip.mp4", { duration: 10, file: "clip.mp4" }) + store.addClip({ name: "clip.mp4", trackId: "v1", start: 0 }) + const rec = store.requestExport("mp4", "1080p", EXPORT_DIR) + expect(rec.file.startsWith(join(EXPORT_DIR, "untitled-project"))).toBe(true) + }) + + test("a render keeps the timeline it was queued from", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + const before = s.exportSnapshot(rec.id) + s.addClip({ name: "clip.mp4", trackId: "v1", start: 10 }) + expect(s.get().clips).toHaveLength(2) + // The queued render is still of the one-clip timeline the user approved. + expect(s.exportSnapshot(rec.id).clips).toHaveLength(1) + expect(before.duration).toBe(10) + }) + + test("an abandoned render can be reclaimed once its lease expires", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + const claimed = s.claimExport(rec.id)! + expect(s.claimExport(rec.id)).toBeNull() + + expect(isClaimable(claimed, Date.parse(claimed.heartbeatAt!) + 1000)).toBe(false) + expect(isClaimable(claimed, Date.parse(claimed.heartbeatAt!) + RENDER_LEASE_MS + 1)).toBe(true) + }) + + test("progress keeps the lease alive", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + const claimed = s.claimExport(rec.id)! + const beat = s.setExportProgress(rec.id, 0.5) + expect(Date.parse(beat.heartbeatAt!)).toBeGreaterThanOrEqual(Date.parse(claimed.heartbeatAt!)) + }) + + test("a finished render is terminal, not claimable", () => { + const s = ready() + const rec = s.requestExport("mp4", "1080p", EXPORT_DIR) + s.claimExport(rec.id) + const done = s.completeExport(rec.id, uploaded(new Uint8Array([1, 2, 3]))) + expect(isClaimable(done, Date.now() + RENDER_LEASE_MS * 10)).toBe(false) + }) +}) + +describe("parseRange", () => { + test("a suffix range returns the end of the file, not the start", () => { + expect(parseRange("bytes=-500", 5000)).toEqual({ start: 4500, end: 4999 }) + }) + + test("an open-ended range runs to the last byte", () => { + expect(parseRange("bytes=100-", 5000)).toEqual({ start: 100, end: 4999 }) + }) + + test("a closed range is clamped to the file", () => { + expect(parseRange("bytes=100-99999", 5000)).toEqual({ start: 100, end: 4999 }) + expect(parseRange("bytes=0-99", 5000)).toEqual({ start: 0, end: 99 }) + }) + + test("a suffix longer than the file returns the whole file", () => { + expect(parseRange("bytes=-99999", 5000)).toEqual({ start: 0, end: 4999 }) + }) + + test("nonsense and out-of-bounds ranges are rejected", () => { + expect(parseRange(undefined, 5000)).toBeNull() + expect(parseRange("items=0-10", 5000)).toBeNull() + expect(parseRange("bytes=5000-", 5000)).toBe("unsatisfiable") + expect(parseRange("bytes=-0", 5000)).toBe("unsatisfiable") + expect(parseRange("bytes=-", 5000)).toBe("unsatisfiable") + }) +}) + +describe("safeMediaName", () => { + test("strips any path from a name", () => { + expect(safeMediaName("../../etc/passwd")).toBe("passwd") + expect(safeMediaName("clip.mp4")).toBe("clip.mp4") + }) + + test("rejects names that are not usable files", () => { + expect(() => safeMediaName("")).toThrow() + expect(() => safeMediaName("..")).toThrow() + expect(() => safeMediaName(".hidden")).toThrow() + }) +}) diff --git a/apps/web/package.json b/apps/web/package.json index af6188d..f89871f 100644 --- a/apps/web/package.json +++ b/apps/web/package.json @@ -33,6 +33,7 @@ "embla-carousel-react": "^8.6.0", "input-otp": "^1.4.2", "lucide-react": "^1.14.0", + "mediabunny": "^1.55.2", "next-themes": "^0.4.6", "radix-ui": "^1.4.3", "react": "^19.2.0", diff --git a/apps/web/src/components/editor/data.ts b/apps/web/src/components/editor/data.ts index 049bbab..91c5210 100644 --- a/apps/web/src/components/editor/data.ts +++ b/apps/web/src/components/editor/data.ts @@ -21,6 +21,46 @@ export type Track = { kind: ClipKind } +export type Segment = { start: number; end: number; text?: string } + +/** What the agent knows about one source file. Mirrors MediaInfo in apps/agent. */ +export type MediaInfo = { + duration: number + silences?: Segment[] + transcript?: Segment[] + bpm?: number + /** Present once real bytes are on disk. Without it a clip cannot be drawn or rendered. */ + file?: string + width?: number + height?: number + fps?: number + hasAudio?: boolean + sizeBytes?: number + peaks?: number[] + analyzedAt?: string +} + +export type ExportStatus = "pending" | "rendering" | "done" | "failed" + +export type ExportRecord = { + id: string + format: string + resolution: string + width: number + height: number + fps: number + createdAt: string + durationSeconds: number + file: string + status: ExportStatus + progress?: number + sizeBytes?: number + completedAt?: string + error?: string + claimedAt?: string + heartbeatAt?: string +} + export type Project = { name: string fps: number @@ -28,7 +68,8 @@ export type Project = { duration: number tracks: Track[] clips: Clip[] - exports?: { id: string; format: string; resolution: string; createdAt: string; file: string }[] + media?: Record + exports?: ExportRecord[] } export const PROJECT: Project = { diff --git a/apps/web/src/components/editor/exports.ts b/apps/web/src/components/editor/exports.ts new file mode 100644 index 0000000..54e3c20 --- /dev/null +++ b/apps/web/src/components/editor/exports.ts @@ -0,0 +1,15 @@ +import type { ExportRecord } from "./data" + +/** + * How long a claimed render may go silent before another editor may take it over. Must match + * RENDER_LEASE_MS in apps/agent; the server is the authority, this only avoids pointless claims. + */ +export const RENDER_LEASE_MS = 60_000 + +/** Queued, or claimed by a worker that has gone quiet: either way, free to pick up. */ +export function isClaimable(rec: ExportRecord, now = Date.now()): boolean { + if (rec.status === "pending") return true + if (rec.status !== "rendering") return false + const last = Date.parse(rec.heartbeatAt ?? rec.claimedAt ?? rec.createdAt) + return Number.isFinite(last) && now - last > RENDER_LEASE_MS +} diff --git a/apps/web/src/components/editor/preview.tsx b/apps/web/src/components/editor/preview.tsx index dd4ed9a..3ed746e 100644 --- a/apps/web/src/components/editor/preview.tsx +++ b/apps/web/src/components/editor/preview.tsx @@ -1,7 +1,8 @@ -import { Pause, Play, SkipBack, SkipForward, Volume2 } from "lucide-react" +import { Pause, Play, SkipBack, SkipForward, Volume2, VolumeX } from "lucide-react" import { Button } from "#/components/ui/button" import { cn } from "#/lib/utils" import { clipsAt, formatTimecode, type Project } from "./data" +import { PREVIEW_HEIGHT, PREVIEW_WIDTH, usePreview } from "./use-preview" type Props = { project: Project @@ -12,36 +13,48 @@ type Props = { className?: string } -const SCENES: Record = { - "intro.mp4": "linear-gradient(135deg, #2b2440 0%, #17161c 60%, #0e0e10 100%)", - "b-roll.mp4": "linear-gradient(160deg, #1d3532 0%, #121a1a 55%, #0e0e10 100%)", - "talking-head.mp4": "linear-gradient(150deg, #3a2a22 0%, #1b1613 55%, #0e0e10 100%)", -} - export function Preview({ project, time, playing, onToggle, onSeek, className }: Props) { + const { canvasRef, error, decoding, muted, toggleMuted, missing } = usePreview({ project, time, playing }) const active = clipsAt(project, time) const video = active.find((c) => c.kind === "video") - const text = active.find((c) => c.kind === "text") + const empty = project.clips.length === 0 return (
-
- {video ? ( +
+ + + {video && ( {video.name} - ) : ( - - Nothing on V1 at this time + )} + {empty && ( + + Nothing on the timeline. Import footage from the Media panel, then ask the assistant to cut it. + + )} + {!empty && missing.length > 0 && ( + + No media on disk for {missing.slice(0, 3).join(", ")} + {missing.length > 3 ? ` and ${missing.length - 3} more` : ""}. Import the real files to see and render them. + + )} + {decoding && ( + + decoding )} - {text && ( - - {text.name} + {error && ( + + {error} )} @@ -71,8 +84,15 @@ export function Preview({ project, time, playing, onToggle, onSeek, className }: / {formatTimecode(project.duration, project.fps)} 16:9 · 1080p · {project.fps} fps -
diff --git a/apps/web/src/components/editor/side-panel.tsx b/apps/web/src/components/editor/side-panel.tsx index 5f14324..6e002fa 100644 --- a/apps/web/src/components/editor/side-panel.tsx +++ b/apps/web/src/components/editor/side-panel.tsx @@ -1,9 +1,10 @@ -import { AudioLines, Captions, Film, Mic, Music, Scissors, Sparkles, Type, Upload, Wand2 } from "lucide-react" -import { useState } from "react" +import { AlertCircle, AudioLines, Captions, Film, Mic, Music, Scissors, Sparkles, Type, Upload, Wand2 } from "lucide-react" +import { useRef, useState } from "react" import { Button } from "#/components/ui/button" import { ScrollArea } from "#/components/ui/scroll-area" import { cn } from "#/lib/utils" -import { CLIP_TONE, type Clip } from "./data" +import { CLIP_TONE, type Clip, type MediaInfo } from "./data" +import type { ImportState } from "./use-media-import" type PanelId = "media" | "ai" | "text" | "audio" @@ -23,11 +24,14 @@ const AI_ACTIONS = [ type Props = { clips: Clip[] + media: Record + imports: ImportState[] + onImport: (files: File[]) => void onSuggest: (prompt: string) => void className?: string } -export function SidePanel({ clips, onSuggest, className }: Props) { +export function SidePanel({ clips, media, imports, onImport, onSuggest, className }: Props) { const [panel, setPanel] = useState("media") return ( @@ -52,7 +56,7 @@ export function SidePanel({ clips, onSuggest, className }: Props) {
- {panel === "media" && } + {panel === "media" && } {panel === "ai" && } {panel === "text" && } {panel === "audio" && } @@ -71,13 +75,26 @@ function PanelTitle({ children, action }: { children: React.ReactNode; action?: ) } -function MediaPanel({ clips }: { clips: Clip[] }) { - const media = clips.filter((c) => c.kind !== "text") +function MediaPanel({ + clips, + media, + imports, + onImport, +}: { + clips: Clip[] + media: Record + imports: ImportState[] + onImport: (files: File[]) => void +}) { + const input = useRef(null) + const entries = Object.entries(media) + const usage = (name: string) => clips.filter((c) => c.name === name).length + return (
+ @@ -85,29 +102,72 @@ function MediaPanel({ clips }: { clips: Clip[] }) { > Media -
    - {media.map((c) => ( -
  • - + {!info.file && ( + + + + )} +
))} -

Drag a file here or onto the timeline to add it.

+ +

+ {entries.length === 0 + ? "No media yet. Import a video or audio file to start editing." + : "Drop files anywhere in the editor to import more."} +

) } diff --git a/apps/web/src/components/editor/timeline.tsx b/apps/web/src/components/editor/timeline.tsx index 108199d..2e04a18 100644 --- a/apps/web/src/components/editor/timeline.tsx +++ b/apps/web/src/components/editor/timeline.tsx @@ -2,7 +2,7 @@ import { Magnet, Scissors, ZoomIn, ZoomOut } from "lucide-react" import { useEffect, useRef, useState } from "react" import { Button } from "#/components/ui/button" import { cn } from "#/lib/utils" -import { CLIP_TONE, clamp, formatShort, type Project } from "./data" +import { CLIP_TONE, clamp, formatShort, type Clip, type MediaInfo, type Project } from "./data" type Props = { project: Project @@ -171,7 +171,7 @@ export function Timeline({ project, time, playing, selectedId, onSelect, onSeek, }} > {c.name} - {c.kind === "audio" && } + {c.kind === "audio" && } ) })} @@ -195,18 +195,38 @@ export function Timeline({ project, time, playing, selectedId, onSelect, onSeek, ) } -function Waveform({ seed }: { seed: string }) { +function Waveform({ clip, media }: { clip: Clip; media?: MediaInfo }) { const bars = 120 - let x = seed.charCodeAt(seed.length - 1) * 9301 + 49297 - const heights = Array.from({ length: bars }, () => { - x = (x * 9301 + 49297) % 233280 - return 0.2 + (x / 233280) * 0.8 - }) + const heights = media?.peaks?.length ? clipPeaks(media.peaks, clip, media.duration, bars) : placeholderPeaks(clip.id, bars) return ( {heights.map((h, i) => ( - + ))} ) } + +/** The measured envelope, windowed to the part of the source this clip actually uses. */ +function clipPeaks(peaks: number[], clip: Clip, mediaDuration: number, bars: number): number[] { + if (mediaDuration <= 0) return placeholderPeaks(clip.id, bars) + const from = ((clip.sourceOffset ?? 0) / mediaDuration) * peaks.length + const to = (((clip.sourceOffset ?? 0) + clip.duration) / mediaDuration) * peaks.length + const span = Math.max(1, to - from) + return Array.from({ length: bars }, (_, i) => { + const start = Math.floor(from + (i / bars) * span) + const end = Math.max(start + 1, Math.floor(from + ((i + 1) / bars) * span)) + let peak = 0 + for (let j = start; j < end && j < peaks.length; j++) peak = Math.max(peak, peaks[j] ?? 0) + return peak + }) +} + +/** Stable stand-in for media that has not been analyzed, so clips still read as audio. */ +function placeholderPeaks(seed: string, bars: number): number[] { + let x = seed.charCodeAt(seed.length - 1) * 9301 + 49297 + return Array.from({ length: bars }, () => { + x = (x * 9301 + 49297) % 233280 + return 0.2 + (x / 233280) * 0.8 + }) +} diff --git a/apps/web/src/components/editor/top-bar.tsx b/apps/web/src/components/editor/top-bar.tsx index 6b4ac31..64d2c9a 100644 --- a/apps/web/src/components/editor/top-bar.tsx +++ b/apps/web/src/components/editor/top-bar.tsx @@ -3,7 +3,9 @@ import { Button } from "#/components/ui/button" import { Separator } from "#/components/ui/separator" import { cn } from "#/lib/utils" import { Tooltip, TooltipContent, TooltipTrigger } from "#/components/ui/tooltip" -import { formatTimecode } from "./data" +import { AGENT_URL } from "#/lib/agent" +import { formatTimecode, type ExportRecord } from "./data" +import type { RenderState } from "./use-render-worker" type Props = { projectName: string @@ -11,9 +13,13 @@ type Props = { fps: number live: boolean lastChange: string | null + render: RenderState + lastExport?: ExportRecord + onExport: () => void + canExport: boolean } -export function TopBar({ projectName, time, fps, live, lastChange }: Props) { +export function TopBar({ projectName, time, fps, live, lastChange, render, lastExport, onExport, canExport }: Props) { return (
@@ -39,18 +45,25 @@ export function TopBar({ projectName, time, fps, live, lastChange }: Props) { {formatTimecode(time, fps)} - - + ) : null} +
) } +function formatBytes(bytes?: number) { + if (!bytes) return "Download" + return bytes >= 1e6 ? `${(bytes / 1e6).toFixed(1)} MB` : `${Math.round(bytes / 1e3)} KB` +} + function IconButton({ label, icon }: { label: string; icon: React.ReactNode }) { return ( diff --git a/apps/web/src/components/editor/use-media-import.ts b/apps/web/src/components/editor/use-media-import.ts new file mode 100644 index 0000000..de89bc8 --- /dev/null +++ b/apps/web/src/components/editor/use-media-import.ts @@ -0,0 +1,63 @@ +import { useCallback, useState } from "react" +import { decodeAudioFile, detectSilences, estimateTempo, forgetAudio, peakEnvelope } from "#/engine/audio" +import { probeFile, uploadMedia } from "#/engine/media" +import { videoCache } from "#/engine/video-cache" +import { agentJson } from "#/lib/agent" + +export type ImportState = { name: string; stage: "probing" | "uploading" | "analyzing" | "done" | "failed"; error?: string } + +/** + * Import real footage: measure it, hand the bytes to the agent, then measure its audio so the + * agent's silence, waveform and tempo answers come from the file rather than from a fixture. + */ +export function useMediaImport() { + const [imports, setImports] = useState([]) + + const update = useCallback((name: string, patch: Partial) => { + setImports((list) => list.map((i) => (i.name === name ? { ...i, ...patch } : i))) + }, []) + + const importFiles = useCallback( + async (files: File[]) => { + if (files.length === 0) return + setImports((list) => [...list.filter((i) => !files.some((f) => f.name === i.name)), ...files.map((f) => ({ name: f.name, stage: "probing" as const }))]) + + for (const file of files) { + try { + const probe = await probeFile(file) + if (!probe.canDecode) throw new Error(`This browser cannot decode ${probe.codec ?? "that codec"}.`) + + update(file.name, { stage: "uploading" }) + const { name } = await uploadMedia(file, probe) + // The bytes changed under this name, so anything decoded from the old ones is stale. + videoCache.clear(name) + forgetAudio(name) + + if (probe.hasAudio) { + update(file.name, { stage: "analyzing" }) + const buffer = await decodeAudioFile(file) + if (buffer) { + await agentJson(`/media/${encodeURIComponent(name)}/analysis`, { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ + silences: detectSilences(buffer), + peaks: peakEnvelope(buffer), + bpm: estimateTempo(buffer) ?? undefined, + }), + }) + } + } + update(file.name, { stage: "done" }) + } catch (err) { + update(file.name, { stage: "failed", error: err instanceof Error ? err.message : String(err) }) + } + } + }, + [update], + ) + + const dismiss = useCallback((name: string) => setImports((list) => list.filter((i) => i.name !== name)), []) + + return { imports, importFiles, dismiss } +} diff --git a/apps/web/src/components/editor/use-preview.ts b/apps/web/src/components/editor/use-preview.ts new file mode 100644 index 0000000..8fde2bf --- /dev/null +++ b/apps/web/src/components/editor/use-preview.ts @@ -0,0 +1,176 @@ +import { useCallback, useEffect, useRef, useState } from "react" +import { audioContext, mixTimeline } from "#/engine/audio" +import { clipsAtTime, drawTimelineFrame, sourceTimeFor, type FrameSources } from "#/engine/compositor" +import { videoCache } from "#/engine/video-cache" +import { fireAndForget } from "#/lib/async" +import type { Project } from "./data" + +/** Preview resolution. Lower than the export on purpose: scrubbing should stay responsive. */ +export const PREVIEW_WIDTH = 1280 +export const PREVIEW_HEIGHT = 720 + +/** Past this much drift the audio node is restarted rather than left to run away. */ +const RESYNC_SECONDS = 0.3 + +/** Only the parts of the project that change the mixdown. */ +function audioSignature(project: Project) { + return project.clips + .filter((c) => c.kind !== "text") + .map((c) => `${c.id}:${c.name}:${c.start}:${c.duration}:${c.sourceOffset ?? 0}:${c.volume ?? 100}`) + .join("|") +} + +/** Media a clip needs but the agent does not have bytes for. Nothing can be drawn for these. */ +export function missingMedia(project: Project): string[] { + const names = new Set() + for (const clip of project.clips) { + if (clip.kind === "text") continue + if (!project.media?.[clip.name]?.file) names.add(clip.name) + } + return [...names] +} + +/** + * Drives the preview: real decoded frames onto a canvas, and the real timeline mixdown + * through WebAudio. Both come from the same engine the exporter uses, so what you watch is + * what gets encoded. + */ +export function usePreview({ project, time, playing }: { project: Project; time: number; playing: boolean }) { + const canvasRef = useRef(null) + const [error, setError] = useState(null) + const [decoding, setDecoding] = useState(false) + + // Draw requests are serialised: the decoder is single-threaded per media and the newest + // requested time is the only one worth painting. + const drawing = useRef(false) + const queued = useRef(null) + + const draw = useCallback( + async (at: number) => { + const canvas = canvasRef.current + const ctx = canvas?.getContext("2d") + if (!canvas || !ctx) return + if (drawing.current) { + queued.current = at + return + } + drawing.current = true + try { + const frames: FrameSources = new Map() + const videoClips = clipsAtTime(project, at).filter((c) => c.kind === "video" && project.media?.[c.name]?.file) + if (videoClips.length > 0) setDecoding(true) + for (const clip of videoClips) { + try { + const frame = await videoCache.getFrameAt(clip.name, sourceTimeFor(clip, at)) + if (frame) frames.set(clip.id, frame.canvas) + } catch (err) { + setError(err instanceof Error ? err.message : String(err)) + } + } + drawTimelineFrame(ctx, { project, time: at, frames, width: canvas.width, height: canvas.height }) + } finally { + drawing.current = false + setDecoding(false) + const next = queued.current + queued.current = null + if (next !== null && next !== at) fireAndForget(draw(next)) + } + }, + [project], + ) + + useEffect(() => { + fireAndForget(draw(time)) + }, [draw, time]) + + // ---- audio --------------------------------------------------------------- + + const mix = useRef<{ signature: string; buffer: AudioBuffer | null } | null>(null) + const node = useRef(null) + /** Context time when playback started, and the timeline time it started from. */ + const anchor = useRef<{ contextTime: number; timelineTime: number } | null>(null) + const [muted, setMuted] = useState(false) + const signature = audioSignature(project) + + /** + * The live project, read inside startAudio without being a dependency of it. + * + * Every agent edit and every render progress report broadcasts a new project object. If the + * callback depended on that object, playback would stop and restart a few times a second + * during a render, which is audible. + */ + const projectRef = useRef(project) + projectRef.current = project + + /** Bumped whenever playback should stop; a slow mixdown checks it before it starts a node. */ + const generation = useRef(0) + + const stopAudio = useCallback(() => { + generation.current += 1 + if (node.current) { + try { + node.current.stop() + } catch { + // already stopped + } + node.current.disconnect() + node.current = null + } + anchor.current = null + }, []) + + const startAudio = useCallback( + async (from: number) => { + stopAudio() + const mine = generation.current + if (muted || projectRef.current.duration <= 0) return + const ctx = audioContext() + if (mix.current?.signature !== signature) { + mix.current = { signature, buffer: await mixTimeline(projectRef.current).catch(() => null) } + } + const buffer = mix.current?.buffer + if (!buffer || from >= buffer.duration) return + if (ctx.state === "suspended") await ctx.resume() + // Mixing and resuming are both awaits, and the user may have paused across either of + // them. Without this the preview sits paused while audio plays on. + if (generation.current !== mine) return + const source = ctx.createBufferSource() + source.buffer = buffer + source.connect(ctx.destination) + source.start(0, from) + node.current = source + anchor.current = { contextTime: ctx.currentTime, timelineTime: from } + }, + [muted, signature, stopAudio], + ) + + const timeRef = useRef(time) + timeRef.current = time + + useEffect(() => { + if (playing) fireAndForget(startAudio(timeRef.current)) + else stopAudio() + return stopAudio + }, [playing, startAudio, stopAudio]) + + // A scrub during playback, or an edit landing mid-play, leaves the audio where it was. + useEffect(() => { + if (!playing || !anchor.current) return + const ctx = audioContext() + const expected = anchor.current.timelineTime + (ctx.currentTime - anchor.current.contextTime) + if (Math.abs(expected - time) > RESYNC_SECONDS) fireAndForget(startAudio(time)) + }, [playing, time, startAudio]) + + const toggleMuted = useCallback(() => { + setMuted((m) => { + if (!m) stopAudio() + return !m + }) + }, [stopAudio]) + + useEffect(() => { + if (playing && !muted && !node.current) fireAndForget(startAudio(timeRef.current)) + }, [muted, playing, startAudio]) + + return { canvasRef, error, decoding, muted, toggleMuted, missing: missingMedia(project) } +} diff --git a/apps/web/src/components/editor/use-project.ts b/apps/web/src/components/editor/use-project.ts index 718c850..3a1aa4b 100644 --- a/apps/web/src/components/editor/use-project.ts +++ b/apps/web/src/components/editor/use-project.ts @@ -1,7 +1,8 @@ import { useEffect, useState } from "react" +import { AGENT_URL } from "#/lib/agent" import { PROJECT, type Project } from "./data" -export const AGENT_URL = (import.meta.env.VITE_EDITAI_AGENT_URL as string | undefined) ?? "http://localhost:8941" +export { AGENT_URL } type State = { project: Project; connected: boolean; revision: number; lastChange: string | null } diff --git a/apps/web/src/components/editor/use-render-worker.ts b/apps/web/src/components/editor/use-render-worker.ts new file mode 100644 index 0000000..df6f08c --- /dev/null +++ b/apps/web/src/components/editor/use-render-worker.ts @@ -0,0 +1,101 @@ +import { useCallback, useEffect, useRef, useState } from "react" +import { renderTimeline, type ExportFormat, type RenderSink } from "#/engine/exporter" +import { AGENT_URL, agentJson } from "#/lib/agent" +import { fireAndForget } from "#/lib/async" +import { isClaimable } from "./exports" +import type { ExportRecord, Project } from "./data" + +export type RenderState = { id: string; progress: number; resolution: string; format: string } | null + +/** Progress is posted through the project store, so report it sparingly. */ +const PROGRESS_STEP = 0.02 + +/** + * Jobs this page has already tried, shared across every instance of the hook. A ref would be + * per-mount, and React remounts effects in development, which is enough to claim twice. + */ +const attempted = new Set() + +/** + * The agent queues renders; the editor performs them. The decoders and the encoder both live + * here, so this is the only place that can turn a timeline into a file. + */ +export function useRenderWorker(project: Project) { + const [state, setState] = useState(null) + const [error, setError] = useState(null) + const busy = useRef(false) + + const run = useCallback(async (job: ExportRecord) => { + busy.current = true + let lastPosted = -1 + try { + // The server hands a job to exactly one claimer, so a second editor open on the same + // project renders nothing rather than encoding a duplicate over the top of this one. + const claim = await fetch(`${AGENT_URL}/exports/${job.id}/claim`, { method: "POST" }) + if (claim.status === 409) return + if (!claim.ok) throw new Error(`Could not claim ${job.id}: ${claim.status}`) + + setError(null) + setState({ id: job.id, progress: 0, resolution: job.resolution, format: job.format }) + + // The timeline as it was when the render was approved. Rendering the live project would + // silently encode edits made after the user agreed to the export. + const { project: snapshot } = await agentJson<{ project: Project }>(`/exports/${job.id}/project`) + + const sink: RenderSink = async ({ data, position }) => { + const res = await fetch(`${AGENT_URL}/exports/${job.id}/chunk?position=${position}`, { + method: "POST", + body: data as BodyInit, + headers: { "content-type": "application/octet-stream" }, + }) + if (!res.ok) throw new Error(`Chunk upload failed with ${res.status}: ${(await res.text()).slice(0, 200)}`) + } + + await renderTimeline({ + project: snapshot, + width: job.width, + height: job.height, + fps: job.fps, + format: job.format === "webm" ? ("webm" as ExportFormat) : ("mp4" as ExportFormat), + sink, + onProgress: (fraction) => { + setState((s) => (s && s.id === job.id ? { ...s, progress: fraction } : s)) + if (fraction - lastPosted < PROGRESS_STEP) return + lastPosted = fraction + fireAndForget( + agentJson(`/exports/${job.id}/progress`, { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ progress: fraction }), + }), + ) + }, + }) + + await agentJson(`/exports/${job.id}/finish`, { method: "POST" }) + } catch (err) { + const message = err instanceof Error ? err.message : String(err) + setError(message) + fireAndForget( + agentJson(`/exports/${job.id}/failed`, { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ error: message }), + }), + ) + } finally { + busy.current = false + setState(null) + } + }, []) + + useEffect(() => { + if (busy.current) return + const job = project.exports?.find((e) => isClaimable(e) && !attempted.has(e.id)) + if (!job) return + attempted.add(job.id) + fireAndForget(run(job)) + }, [project, run]) + + return { state, error } +} diff --git a/apps/web/src/engine/audio.test.ts b/apps/web/src/engine/audio.test.ts new file mode 100644 index 0000000..c0ff5c6 --- /dev/null +++ b/apps/web/src/engine/audio.test.ts @@ -0,0 +1,93 @@ +import { describe, expect, test } from "vitest" +import { detectSilences, estimateTempo, peakEnvelope } from "./audio" + +const SAMPLE_RATE = 48000 + +/** + * The analyzers only read `sampleRate` and channel 0, so a plain Float32Array standing in for + * an AudioBuffer exercises the real code without a WebAudio implementation in the test env. + */ +function buffer(samples: Float32Array): AudioBuffer { + return { sampleRate: SAMPLE_RATE, length: samples.length, duration: samples.length / SAMPLE_RATE, numberOfChannels: 1, getChannelData: () => samples } as unknown as AudioBuffer +} + +/** A tone with silent gaps, the same shape as the sample voiceover. */ +function tone(durationSeconds: number, gaps: { start: number; end: number }[]): AudioBuffer { + const samples = new Float32Array(Math.round(durationSeconds * SAMPLE_RATE)) + for (let i = 0; i < samples.length; i++) { + const t = i / SAMPLE_RATE + samples[i] = gaps.some((g) => t >= g.start && t < g.end) ? 0 : 0.4 * Math.sin(2 * Math.PI * 220 * t) + } + return buffer(samples) +} + +describe("detectSilences", () => { + test("finds the gaps that are really there", () => { + const gaps = [ + { start: 1, end: 2 }, + { start: 4.5, end: 5.4 }, + ] + const found = detectSilences(tone(8, gaps)) + expect(found).toHaveLength(2) + expect(found[0]!.start).toBeCloseTo(1, 1) + expect(found[0]!.end).toBeCloseTo(2, 1) + expect(found[1]!.start).toBeCloseTo(4.5, 1) + }) + + test("ignores gaps shorter than the minimum", () => { + expect(detectSilences(tone(4, [{ start: 1, end: 1.2 }]), { minDuration: 0.5 })).toEqual([]) + }) + + test("room tone below the threshold still counts as silence", () => { + const samples = new Float32Array(SAMPLE_RATE * 3) + for (let i = 0; i < samples.length; i++) { + const t = i / SAMPLE_RATE + // -60 dBFS hiss between 1s and 2s: quiet, but nowhere near zero. + samples[i] = t >= 1 && t < 2 ? 0.001 * Math.sin(2 * Math.PI * 3000 * t) : 0.4 * Math.sin(2 * Math.PI * 220 * t) + } + const found = detectSilences(buffer(samples)) + expect(found).toHaveLength(1) + expect(found[0]!.start).toBeCloseTo(1, 1) + }) + + test("audio with no gaps has no silences", () => { + expect(detectSilences(tone(5, []))).toEqual([]) + }) +}) + +describe("peakEnvelope", () => { + test("tracks amplitude across the file", () => { + const samples = new Float32Array(SAMPLE_RATE * 2) + // Quiet first half, loud second half. + for (let i = 0; i < samples.length; i++) samples[i] = (i < samples.length / 2 ? 0.1 : 0.9) * Math.sin(i) + const peaks = peakEnvelope(buffer(samples), 10) + expect(peaks).toHaveLength(10) + expect(peaks[0]!).toBeLessThan(0.2) + expect(peaks[9]!).toBeGreaterThan(0.8) + expect(Math.max(...peaks)).toBeLessThanOrEqual(1) + }) +}) + +describe("estimateTempo", () => { + test("recovers the tempo of a click track", () => { + const seconds = 16 + const bpm = 120 + const samples = new Float32Array(SAMPLE_RATE * seconds) + for (let i = 0; i < samples.length; i++) { + const t = i / SAMPLE_RATE + const sinceBeat = t % (60 / bpm) + samples[i] = 0.9 * Math.sin(2 * Math.PI * 760 * t) * Math.exp(-26 * sinceBeat) + } + // Octave errors are the usual failure of autocorrelation; accept the beat or its double. + const found = estimateTempo(buffer(samples)) + expect([bpm, bpm * 2, bpm / 2]).toContain(found) + }) + + test("silence has no tempo", () => { + expect(estimateTempo(buffer(new Float32Array(SAMPLE_RATE * 4)))).toBeNull() + }) + + test("audio too short to hold a beat returns nothing", () => { + expect(estimateTempo(buffer(new Float32Array(256)))).toBeNull() + }) +}) diff --git a/apps/web/src/engine/audio.ts b/apps/web/src/engine/audio.ts new file mode 100644 index 0000000..788ce20 --- /dev/null +++ b/apps/web/src/engine/audio.ts @@ -0,0 +1,264 @@ +import { ALL_FORMATS, AudioBufferSink, BlobSource, Input } from "mediabunny" +import { mediaUrl } from "#/lib/agent" +import type { Project } from "#/components/editor/data" +import { inputFor } from "./media" + +/** 48 kHz stereo is what the AAC/Opus encoders want anyway. */ +export const EXPORT_SAMPLE_RATE = 48000 + +let sharedContext: AudioContext | null = null + +/** One context for the page. Created suspended; playback resumes it on a user gesture. */ +export function audioContext(): AudioContext { + if (!sharedContext) sharedContext = new AudioContext({ sampleRate: EXPORT_SAMPLE_RATE }) + return sharedContext +} + +const decoded = new Map>() + +/** + * Whole-file audio, decoded once per media. + * + * `decodeAudioData` is the fast path and handles mp4/mp3/wav, but it refuses containers the + * browser will not demux (some .mov, .mkv). mediabunny can demux those, so it decodes the + * track packet by packet and the pieces are stitched back together. + */ +export function decodeAudio(name: string): Promise { + const existing = decoded.get(name) + if (existing) return existing + const promise = decodeAudioUncached(name).catch((err) => { + decoded.delete(name) + throw err + }) + decoded.set(name, promise) + return promise +} + +export function forgetAudio(name?: string) { + if (name) decoded.delete(name) + else decoded.clear() +} + +async function decodeAudioUncached(name: string): Promise { + const ctx = audioContext() + try { + const res = await fetch(mediaUrl(name)) + if (!res.ok) throw new Error(`Could not fetch ${name}: ${res.status}`) + return await ctx.decodeAudioData(await res.arrayBuffer()) + } catch { + return await decodeViaMediabunny(inputFor(name), ctx) + } +} + +/** Same decode for a file the user just picked, before it has been uploaded. */ +export async function decodeAudioFile(file: File): Promise { + const ctx = audioContext() + try { + return await ctx.decodeAudioData(await file.arrayBuffer()) + } catch { + return await decodeViaMediabunny(new Input({ source: new BlobSource(file), formats: ALL_FORMATS }), ctx) + } +} + +async function decodeViaMediabunny(input: Input, ctx: BaseAudioContext): Promise { + const track = await input.getPrimaryAudioTrack() + if (!track || !(await track.canDecode())) return null + const sink = new AudioBufferSink(track) + const chunks: { buffer: AudioBuffer; timestamp: number }[] = [] + let channels = 0 + let sampleRate = 0 + let end = 0 + for await (const wrapped of sink.buffers()) { + chunks.push({ buffer: wrapped.buffer, timestamp: wrapped.timestamp }) + channels = Math.max(channels, wrapped.buffer.numberOfChannels) + sampleRate = sampleRate || wrapped.buffer.sampleRate + end = Math.max(end, wrapped.timestamp + wrapped.buffer.duration) + } + if (chunks.length === 0) return null + const out = ctx.createBuffer(channels, Math.ceil(end * sampleRate), sampleRate) + for (const { buffer, timestamp } of chunks) { + const offset = Math.round(timestamp * sampleRate) + for (let c = 0; c < channels; c++) { + // A mono chunk in a stereo file feeds every output channel. + const source = buffer.getChannelData(Math.min(c, buffer.numberOfChannels - 1)) + const target = out.getChannelData(c) + const count = Math.min(source.length, target.length - offset) + if (count > 0) target.set(source.subarray(0, count), offset) + } + } + return out +} + +/** Clips that contribute audio: the audio tracks, plus video whose source has sound. */ +export function audibleClips(project: Project) { + return project.clips.filter((c) => { + if (c.kind === "text") return false + if ((c.volume ?? 100) <= 0) return false + if (c.kind === "audio") return true + return project.media?.[c.name]?.hasAudio === true + }) +} + +/** + * The whole timeline mixed to one buffer: every clip placed at its start, read from its + * source offset, scaled by its volume. This is both what the preview plays and what the + * exporter muxes, so the two can never drift apart. + */ +export async function mixTimeline(project: Project, duration = project.duration): Promise { + const clips = audibleClips(project) + if (clips.length === 0 || duration <= 0) return null + + const ctx = audioContext() + const sampleRate = EXPORT_SAMPLE_RATE + const length = Math.ceil(duration * sampleRate) + const out = ctx.createBuffer(2, length, sampleRate) + const left = out.getChannelData(0) + const right = out.getChannelData(1) + + const sources = await Promise.all( + clips.map(async (clip) => ({ clip, buffer: await decodeAudio(clip.name).catch(() => null) })), + ) + + for (const { clip, buffer } of sources) { + if (!buffer) continue + const gain = (clip.volume ?? 100) / 100 + const ratio = buffer.sampleRate / sampleRate + const startSample = Math.round(clip.start * sampleRate) + const count = Math.min(Math.round(clip.duration * sampleRate), length - startSample) + if (count <= 0) continue + const sourceStart = (clip.sourceOffset ?? 0) * buffer.sampleRate + const l = buffer.getChannelData(0) + const r = buffer.numberOfChannels > 1 ? buffer.getChannelData(1) : l + + for (let i = 0; i < count; i++) { + // Nearest-sample read. Sources are decoded at their own rate, and every real file we + // accept is 44.1k or 48k, so the error is under a sample period. + const s = Math.round(sourceStart + i * ratio) + if (s < 0 || s >= l.length) continue + left[startSample + i]! += l[s]! * gain + right[startSample + i]! += r[s]! * gain + } + } + + limit(left) + limit(right) + return out +} + +/** Summing tracks can exceed full scale; scale the whole channel back rather than clip it. */ +function limit(channel: Float32Array) { + let peak = 0 + for (let i = 0; i < channel.length; i++) { + const abs = Math.abs(channel[i]!) + if (abs > peak) peak = abs + } + if (peak <= 1) return + const scale = 1 / peak + for (let i = 0; i < channel.length; i++) channel[i]! *= scale +} + +export type SilenceRange = { start: number; end: number } + +/** + * Silent ranges measured from the decoded audio, in source seconds. + * + * Windowed RMS rather than a per-sample threshold: a single zero crossing is not silence, + * and room tone sits well above zero but well below speech. + */ +export function detectSilences( + buffer: AudioBuffer, + { minDuration = 0.5, thresholdDb = -45, windowMs = 20 }: { minDuration?: number; thresholdDb?: number; windowMs?: number } = {}, +): SilenceRange[] { + const threshold = 10 ** (thresholdDb / 20) + const windowSize = Math.max(1, Math.round((windowMs / 1000) * buffer.sampleRate)) + const data = buffer.getChannelData(0) + const ranges: SilenceRange[] = [] + let runStart: number | null = null + + for (let i = 0; i < data.length; i += windowSize) { + const end = Math.min(i + windowSize, data.length) + let sum = 0 + for (let j = i; j < end; j++) sum += data[j]! * data[j]! + const quiet = Math.sqrt(sum / (end - i)) < threshold + if (quiet && runStart === null) runStart = i + if (!quiet && runStart !== null) { + pushRange(ranges, runStart / buffer.sampleRate, i / buffer.sampleRate, minDuration) + runStart = null + } + } + if (runStart !== null) pushRange(ranges, runStart / buffer.sampleRate, data.length / buffer.sampleRate, minDuration) + return ranges +} + +function pushRange(ranges: SilenceRange[], start: number, end: number, minDuration: number) { + if (end - start < minDuration) return + ranges.push({ start: round(start), end: round(end) }) +} + +/** Peak envelope for the timeline waveform: `count` buckets of max amplitude, 0..1. */ +export function peakEnvelope(buffer: AudioBuffer, count = 400): number[] { + const data = buffer.getChannelData(0) + const peaks = new Array(count).fill(0) + const per = data.length / count + for (let i = 0; i < count; i++) { + const start = Math.floor(i * per) + const end = Math.min(Math.floor((i + 1) * per), data.length) + let peak = 0 + for (let j = start; j < end; j++) { + const abs = Math.abs(data[j]!) + if (abs > peak) peak = abs + } + peaks[i] = round(peak) + } + return peaks +} + +const round = (n: number) => Math.round(n * 1000) / 1000 + +/** + * Tempo, estimated from the onset envelope. + * + * Energy is summed in short hops, the rising part of its difference is the onset strength, + * and the lag whose autocorrelation peaks over a musical range of periods is the beat. Good + * enough to cut to; it will not track a song that changes tempo. + */ +export function estimateTempo(buffer: AudioBuffer, { minBpm = 60, maxBpm = 180 }: { minBpm?: number; maxBpm?: number } = {}): number | null { + const hop = Math.round(buffer.sampleRate * 0.01) + const data = buffer.getChannelData(0) + const frames = Math.floor(data.length / hop) + if (frames < 128) return null + + const energy = new Float32Array(frames) + for (let f = 0; f < frames; f++) { + let sum = 0 + for (let i = f * hop; i < (f + 1) * hop; i++) sum += data[i]! * data[i]! + energy[f] = Math.sqrt(sum / hop) + } + + const onset = new Float32Array(frames) + let mean = 0 + for (let f = 1; f < frames; f++) { + onset[f] = Math.max(0, energy[f]! - energy[f - 1]!) + mean += onset[f]! + } + mean /= frames + if (mean <= 0) return null + for (let f = 0; f < frames; f++) onset[f]! -= mean + + const framesPerSecond = buffer.sampleRate / hop + const minLag = Math.floor((60 / maxBpm) * framesPerSecond) + const maxLag = Math.ceil((60 / minBpm) * framesPerSecond) + let bestLag = 0 + let best = 0 + for (let lag = minLag; lag <= maxLag && lag < frames; lag++) { + let sum = 0 + for (let f = lag; f < frames; f++) sum += onset[f]! * onset[f - lag]! + const score = sum / (frames - lag) + if (score > best) { + best = score + bestLag = lag + } + } + if (!bestLag || best <= 0) return null + return Math.round((60 * framesPerSecond) / bestLag) +} diff --git a/apps/web/src/engine/compositor.test.ts b/apps/web/src/engine/compositor.test.ts new file mode 100644 index 0000000..0b3a28b --- /dev/null +++ b/apps/web/src/engine/compositor.test.ts @@ -0,0 +1,67 @@ +import { describe, expect, test } from "vitest" +import type { Project } from "#/components/editor/data" +import { audibleClips } from "./audio" +import { clipsAtTime, sourceTimeFor } from "./compositor" + +const project: Project = { + name: "test", + fps: 30, + duration: 20, + tracks: [ + { id: "v1", label: "V1", kind: "video" }, + { id: "t1", label: "T1", kind: "text" }, + { id: "a1", label: "A1", kind: "audio" }, + ], + clips: [ + { id: "c1", name: "a.mp4", kind: "video", trackId: "v1", start: 0, duration: 5, sourceOffset: 2 }, + { id: "c2", name: "b.mp4", kind: "video", trackId: "v1", start: 5, duration: 5, sourceOffset: 0 }, + { id: "c3", name: "Title", kind: "text", trackId: "t1", start: 1, duration: 3, sourceOffset: 0 }, + { id: "c4", name: "vo.wav", kind: "audio", trackId: "a1", start: 0, duration: 20, sourceOffset: 0, volume: 100 }, + { id: "c5", name: "muted.wav", kind: "audio", trackId: "a1", start: 0, duration: 20, sourceOffset: 0, volume: 0 }, + ], + media: { + "a.mp4": { duration: 30, width: 1920, height: 1080, hasAudio: true, file: "a.mp4" }, + "b.mp4": { duration: 30, width: 1920, height: 1080, hasAudio: false, file: "b.mp4" }, + "vo.wav": { duration: 30, hasAudio: true, file: "vo.wav" }, + "muted.wav": { duration: 30, hasAudio: true, file: "muted.wav" }, + }, +} + +describe("clipsAtTime", () => { + test("a clip is live from its start up to but not including its end", () => { + expect(clipsAtTime(project, 0).map((c) => c.id)).toEqual(["c1", "c4", "c5"]) + expect(clipsAtTime(project, 5).map((c) => c.id)).toContain("c2") + expect(clipsAtTime(project, 5).map((c) => c.id)).not.toContain("c1") + }) + + test("overlapping tracks are all live at once", () => { + expect(clipsAtTime(project, 2).map((c) => c.id).sort()).toEqual(["c1", "c3", "c4", "c5"]) + }) + + test("past the end of the timeline nothing is live", () => { + expect(clipsAtTime(project, 25)).toEqual([]) + }) +}) + +describe("sourceTimeFor", () => { + test("a trimmed clip reads from its source offset", () => { + const clip = project.clips[0]! + expect(sourceTimeFor(clip, 0)).toBe(2) + expect(sourceTimeFor(clip, 3)).toBe(5) + }) + + test("a clip that starts later on the timeline still reads from its own offset", () => { + expect(sourceTimeFor(project.clips[1]!, 7)).toBe(2) + }) +}) + +describe("audibleClips", () => { + test("silent video and muted audio contribute nothing to the mix", () => { + const ids = audibleClips(project).map((c) => c.id) + expect(ids).toContain("c1") // video whose source has audio + expect(ids).toContain("c4") + expect(ids).not.toContain("c2") // video with no audio track + expect(ids).not.toContain("c5") // volume 0 + expect(ids).not.toContain("c3") // text + }) +}) diff --git a/apps/web/src/engine/compositor.ts b/apps/web/src/engine/compositor.ts new file mode 100644 index 0000000..f7ba1d7 --- /dev/null +++ b/apps/web/src/engine/compositor.ts @@ -0,0 +1,109 @@ +import type { Clip, Project } from "#/components/editor/data" + +/** Decoded frames for one instant, keyed by **clip id**: two clips can share one media file. */ +export type FrameSources = Map + +export type Surface = CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D + +/** Clips live on tracks; a clip on a later track paints over one on an earlier track. */ +function byTrackOrder(project: Project) { + const order = new Map(project.tracks.map((t, i) => [t.id, i])) + return (a: Clip, b: Clip) => (order.get(a.trackId) ?? 0) - (order.get(b.trackId) ?? 0) +} + +export const clipsAtTime = (project: Project, time: number) => + project.clips.filter((c) => time >= c.start && time < c.start + c.duration) + +/** Where in the source file a clip is at a given timeline time. */ +export const sourceTimeFor = (clip: Clip, time: number) => (clip.sourceOffset ?? 0) + (time - clip.start) + +/** + * One frame of the timeline, composited onto a 2D surface. + * + * Deliberately Canvas2D rather than the GPU compositor OpenCut uses: EditAI's timeline is + * stacked video, text and audio with no effects, masks or blend modes, and 2D covers all of + * it exactly. The trade is that this cannot grow filters without a real compositor. + */ +export function drawTimelineFrame( + ctx: Surface, + { project, time, frames, width, height }: { project: Project; time: number; frames: FrameSources; width: number; height: number }, +) { + ctx.clearRect(0, 0, width, height) + ctx.fillStyle = "#000000" + ctx.fillRect(0, 0, width, height) + + const active = clipsAtTime(project, time).sort(byTrackOrder(project)) + + for (const clip of active) { + if (clip.kind !== "video") continue + const frame = frames.get(clip.id) + if (frame) drawCover(ctx, frame, width, height) + } + + for (const clip of active) { + if (clip.kind === "text") drawCaption(ctx, clip.name, width, height) + } +} + +/** Fill the frame, cropping the overflowing axis, the way a preview monitor would. */ +function drawCover(ctx: Surface, source: CanvasImageSource, width: number, height: number) { + const sw = sourceWidth(source) + const sh = sourceHeight(source) + if (!sw || !sh) return + const scale = Math.max(width / sw, height / sh) + const w = sw * scale + const h = sh * scale + ctx.drawImage(source, (width - w) / 2, (height - h) / 2, w, h) +} + +/** Each kind of image source reports its intrinsic size under a different name. */ +function sourceWidth(source: CanvasImageSource): number { + if ("videoWidth" in source) return source.videoWidth + if ("naturalWidth" in source) return source.naturalWidth + if ("width" in source) return Number(source.width) + return 0 +} + +function sourceHeight(source: CanvasImageSource): number { + if ("videoHeight" in source) return source.videoHeight + if ("naturalHeight" in source) return source.naturalHeight + if ("height" in source) return Number(source.height) + return 0 +} + +function drawCaption(ctx: Surface, text: string, width: number, height: number) { + const fontSize = Math.round(height * 0.055) + ctx.save() + ctx.font = `600 ${fontSize}px system-ui, -apple-system, "Segoe UI", Inter, sans-serif` + ctx.textAlign = "center" + ctx.textBaseline = "alphabetic" + + const lines = wrap(ctx, text, width * 0.84) + const lineHeight = fontSize * 1.2 + const baseline = height * 0.86 - (lines.length - 1) * lineHeight + + ctx.shadowColor = "rgba(0,0,0,0.7)" + ctx.shadowBlur = fontSize * 0.35 + ctx.shadowOffsetY = fontSize * 0.05 + ctx.fillStyle = "#ffffff" + for (const [i, line] of lines.entries()) ctx.fillText(line, width / 2, baseline + i * lineHeight) + ctx.restore() +} + +/** Greedy wrap. A single word wider than the box is left to overflow rather than broken. */ +function wrap(ctx: Surface, text: string, maxWidth: number): string[] { + const words = text.split(/\s+/).filter(Boolean) + if (words.length === 0) return [] + const lines: string[] = [] + let line = words[0]! + for (const word of words.slice(1)) { + const candidate = `${line} ${word}` + if (ctx.measureText(candidate).width <= maxWidth) line = candidate + else { + lines.push(line) + line = word + } + } + lines.push(line) + return lines +} diff --git a/apps/web/src/engine/exporter.ts b/apps/web/src/engine/exporter.ts new file mode 100644 index 0000000..bd751fc --- /dev/null +++ b/apps/web/src/engine/exporter.ts @@ -0,0 +1,164 @@ +import { + AudioBufferSource, + BufferTarget, + CanvasSource, + Mp4OutputFormat, + Output, + QUALITY_HIGH, + QUALITY_MEDIUM, + QUALITY_VERY_HIGH, + StreamTarget, + WebMOutputFormat, +} from "mediabunny" +import type { Project } from "#/components/editor/data" +import { mixTimeline } from "./audio" +import { clipsAtTime, drawTimelineFrame, sourceTimeFor, type FrameSources } from "./compositor" +import { VideoCache } from "./video-cache" + +export type ExportFormat = "mp4" | "webm" +export type ExportQuality = "medium" | "high" | "very_high" + +const QUALITY = { medium: QUALITY_MEDIUM, high: QUALITY_HIGH, very_high: QUALITY_VERY_HIGH } + +/** 8 MiB slices: few enough requests to be cheap, small enough to stay off the heap. */ +const CHUNK_SIZE = 8 * 1024 * 1024 + +/** Where encoded bytes go as they are produced. Positions are revisited, so this is not an append. */ +export type RenderSink = (chunk: { data: Uint8Array; position: number }) => Promise + +export type RenderOptions = { + project: Project + width: number + height: number + fps: number + format?: ExportFormat + quality?: ExportQuality + onProgress?: (fraction: number) => void + signal?: AbortSignal + /** Given, the encode streams out through it and this returns null instead of a Blob. */ + sink?: RenderSink +} + +function makeCanvas(width: number, height: number): HTMLCanvasElement | OffscreenCanvas { + if (typeof OffscreenCanvas !== "undefined") return new OffscreenCanvas(width, height) + const canvas = document.createElement("canvas") + canvas.width = width + canvas.height = height + return canvas +} + +/** mp4 wants AAC, but not every browser build can encode it; Opus in an mp4 is the fallback. */ +async function pickAudioCodec(format: ExportFormat, buffer: AudioBuffer): Promise<"aac" | "opus"> { + if (format === "webm") return "opus" + if (typeof AudioEncoder === "undefined") return "opus" + const { supported } = await AudioEncoder.isConfigSupported({ + codec: "mp4a.40.2", + sampleRate: buffer.sampleRate, + numberOfChannels: buffer.numberOfChannels, + bitrate: 192_000, + }) + return supported ? "aac" : "opus" +} + +/** + * Encode the timeline to a real video file, in the browser. + * + * This is the same shape as OpenCut's scene exporter: composite each frame onto a canvas, hand + * the canvas to a WebCodecs-backed `CanvasSource`, and mux with mediabunny. The audio mixdown is + * added up front as one buffer rather than per frame, which is what keeps sound in sync with a + * video track whose frame durations are only nominally constant. + */ +export async function renderTimeline({ + project, + width, + height, + fps, + format = "mp4", + quality = "high", + onProgress, + signal, + sink, +}: RenderOptions): Promise { + if (project.duration <= 0) throw new Error("There is nothing on the timeline to render.") + + const canvas = makeCanvas(width, height) + const ctx = canvas.getContext("2d") as CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D | null + if (!ctx) throw new Error("Could not get a 2D context for the render surface.") + + const buffered = sink ? null : new BufferTarget() + const output = new Output({ + format: format === "webm" ? new WebMOutputFormat() : new Mp4OutputFormat(), + target: + buffered ?? + new StreamTarget( + new WritableStream({ + // Awaited, so the muxer applies backpressure instead of queueing chunks in memory. + write: (chunk) => sink!({ data: chunk.data, position: chunk.position }), + }), + { chunked: true, chunkSize: CHUNK_SIZE }, + ), + }) + const videoSource = new CanvasSource(canvas, { + codec: format === "webm" ? "vp9" : "avc", + bitrate: QUALITY[quality], + }) + output.addVideoTrack(videoSource, { frameRate: fps }) + + const audio = await mixTimeline(project) + let audioSource: AudioBufferSource | null = null + if (audio) { + audioSource = new AudioBufferSource({ codec: await pickAudioCodec(format, audio), bitrate: QUALITY[quality] }) + output.addAudioTrack(audioSource) + } + + await output.start() + + // A cache of its own: the preview's decoders are mid-scrub and must not be dragged along. + const cache = new VideoCache() + let finished = false + + try { + if (audioSource && audio) { + await audioSource.add(audio) + audioSource.close() + } + + // Ceil, not round: a timeline that is not a whole number of frames long must still be + // covered to its end, and the final frame is shortened so the video lasts exactly as long + // as the audio rather than up to half a frame more. + const frameCount = Math.max(1, Math.ceil(project.duration * fps - 1e-9)) + for (let i = 0; i < frameCount; i++) { + if (signal?.aborted) throw new DOMException("Render cancelled", "AbortError") + const time = i / fps + const frames: FrameSources = new Map() + for (const clip of clipsAtTime(project, time)) { + if (clip.kind !== "video") continue + const frame = await cache.getFrameAt(clip.name, sourceTimeFor(clip, time)).catch(() => null) + // Keyed by clip, not by media: two clips of one file at different offsets are two frames. + if (frame) frames.set(clip.id, frame.canvas) + } + drawTimelineFrame(ctx, { project, time, frames, width, height }) + await videoSource.add(time, Math.min(1 / fps, project.duration - time)) + onProgress?.(i / frameCount) + } + + videoSource.close() + await output.finalize() + finished = true + } finally { + cache.clear() + if (!finished) { + // An encoder left open after a failure holds hardware and can break the next render. + videoSource.close() + audioSource?.close() + await output.cancel().catch(() => undefined) + } + } + + onProgress?.(1) + if (!buffered) return null + + const bytes = buffered.buffer + if (!bytes) throw new Error("The encoder produced no output.") + return new Blob([bytes], { type: format === "webm" ? "video/webm" : "video/mp4" }) +} diff --git a/apps/web/src/engine/media.ts b/apps/web/src/engine/media.ts new file mode 100644 index 0000000..cf50d1d --- /dev/null +++ b/apps/web/src/engine/media.ts @@ -0,0 +1,71 @@ +import { ALL_FORMATS, BlobSource, Input, UrlSource } from "mediabunny" +import { agentJson, mediaUrl } from "#/lib/agent" + +export type Probe = { + duration: number + width?: number + height?: number + fps?: number + hasAudio: boolean + hasVideo: boolean + /** False when the browser has no decoder for this codec, which export would hit later. */ + canDecode: boolean + codec: string | null +} + +/** + * Inputs are opened over HTTP range requests rather than downloaded, so seeking a large + * file costs a couple of requests instead of the whole file. One per media name, kept for + * the life of the page: a decoder's own caches are what make scrubbing fast. + */ +const inputs = new Map() + +export function inputFor(name: string): Input { + const existing = inputs.get(name) + if (existing) return existing + const input = new Input({ source: new UrlSource(mediaUrl(name)), formats: ALL_FORMATS }) + inputs.set(name, input) + return input +} + +/** Drop a media's decoder, e.g. after it is re-imported under the same name. */ +export function forgetMedia(name: string) { + inputs.delete(name) +} + +/** Measure a file the user picked, before it is uploaded. */ +export async function probeFile(file: File): Promise { + const input = new Input({ source: new BlobSource(file), formats: ALL_FORMATS }) + const [duration, videoTrack, audioTrack] = await Promise.all([ + input.computeDuration(), + input.getPrimaryVideoTrack(), + input.getPrimaryAudioTrack(), + ]) + if (!videoTrack && !audioTrack) throw new Error(`${file.name} has no video or audio track.`) + const canDecode = videoTrack ? await videoTrack.canDecode() : audioTrack ? await audioTrack.canDecode() : false + return { + duration, + width: videoTrack?.displayWidth, + height: videoTrack?.displayHeight, + fps: videoTrack ? ((await videoTrack.computePacketStats(120)).averagePacketRate ?? undefined) : undefined, + hasAudio: Boolean(audioTrack), + hasVideo: Boolean(videoTrack), + canDecode, + codec: videoTrack?.codec ?? audioTrack?.codec ?? null, + } +} + +/** Upload the bytes and register the measurements. The agent stores both. */ +export async function uploadMedia(file: File, probe: Probe): Promise<{ name: string }> { + const params = new URLSearchParams({ duration: String(probe.duration), hasAudio: String(probe.hasAudio) }) + if (probe.width) params.set("width", String(Math.round(probe.width))) + if (probe.height) params.set("height", String(Math.round(probe.height))) + if (probe.fps) params.set("fps", String(Math.round(probe.fps * 1000) / 1000)) + const result = await agentJson<{ name: string }>(`/media/${encodeURIComponent(file.name)}?${params}`, { + method: "POST", + body: file, + headers: { "content-type": file.type || "application/octet-stream" }, + }) + forgetMedia(result.name) + return result +} diff --git a/apps/web/src/engine/video-cache.ts b/apps/web/src/engine/video-cache.ts new file mode 100644 index 0000000..13ab43e --- /dev/null +++ b/apps/web/src/engine/video-cache.ts @@ -0,0 +1,196 @@ +import { CanvasSink, type WrappedCanvas } from "mediabunny" +import { fireAndForget } from "#/lib/async" +import { inputFor } from "./media" + +type SinkData = { + sink: CanvasSink + iterator: AsyncGenerator | null + currentFrame: WrappedCanvas | null + nextFrame: WrappedCanvas | null + lastTime: number + prefetching: boolean + prefetchPromise: Promise | null +} + +/** Seeking is expensive; playing forward is not. Within this many seconds ahead, decode through. */ +const ITERATE_AHEAD_SECONDS = 2 + +/** + * Decoded frames for the preview, one decoder per media file. + * + * Ported from OpenCut's video cache. The shape matters: a naive `getCanvas(t)` per frame + * re-seeks the decoder every time and plays back at a few frames a second. Instead a forward + * iterator is kept open and the next frame is decoded ahead of being asked for, so ordinary + * playback never seeks, and a scrub falls back to a real seek only when the target is behind + * the head or too far in front of it. + */ +export class VideoCache { + private sinks = new Map() + private initPromises = new Map>() + /** Frame requests for one media must not interleave: they share a single decoder. */ + private frameChain = new Map>() + private seekGenerations = new Map() + + async getFrameAt(name: string, time: number): Promise { + await this.ensureSink(name) + const sinkData = this.sinks.get(name) + if (!sinkData) return null + + // A newer request supersedes an older one: the old one resolves with whatever is current + // rather than dragging the decoder back to a timestamp nobody is looking at any more. + const generation = (this.seekGenerations.get(name) ?? 0) + 1 + this.seekGenerations.set(name, generation) + + const previous = this.frameChain.get(name) ?? Promise.resolve() + const current = previous.then(() => { + if (this.seekGenerations.get(name) !== generation) return sinkData.currentFrame + return this.resolveFrame(sinkData, time) + }) + this.frameChain.set( + name, + current.catch(() => {}), + ) + return current + } + + clear(name?: string) { + const names = name ? [name] : [...this.sinks.keys()] + for (const key of names) { + fireAndForget(this.sinks.get(key)?.iterator?.return()) + this.sinks.delete(key) + this.initPromises.delete(key) + this.frameChain.delete(key) + this.seekGenerations.delete(key) + } + } + + private async resolveFrame(sinkData: SinkData, time: number): Promise { + if (sinkData.nextFrame && sinkData.nextFrame.timestamp <= time) { + sinkData.currentFrame = sinkData.nextFrame + sinkData.nextFrame = null + this.startPrefetch(sinkData) + } + + if (sinkData.currentFrame && isFrameValid(sinkData.currentFrame, time)) { + this.startPrefetch(sinkData) + return sinkData.currentFrame + } + + if (sinkData.iterator && sinkData.currentFrame && time >= sinkData.lastTime && time < sinkData.lastTime + ITERATE_AHEAD_SECONDS) { + const frame = await this.iterateToTime(sinkData, time) + if (frame) { + this.startPrefetch(sinkData) + return frame + } + } + + const frame = await this.seekToTime(sinkData, time) + if (frame) this.startPrefetch(sinkData) + return frame + } + + private async iterateToTime(sinkData: SinkData, targetTime: number): Promise { + if (!sinkData.iterator) return null + try { + while (true) { + if (sinkData.prefetching && sinkData.prefetchPromise) await sinkData.prefetchPromise + + if (sinkData.nextFrame && sinkData.nextFrame.timestamp <= targetTime + 0.05) { + sinkData.currentFrame = sinkData.nextFrame + sinkData.nextFrame = null + } else { + const { value: frame, done } = await sinkData.iterator.next() + if (done || !frame) break + sinkData.currentFrame = frame + } + + const frame = sinkData.currentFrame + if (!frame) break + sinkData.lastTime = frame.timestamp + if (isFrameValid(frame, targetTime)) return frame + if (frame.timestamp > targetTime + 1) break + } + } catch (error) { + console.warn("[editai] frame iterator failed, will re-seek:", error) + sinkData.iterator = null + } + return null + } + + private async seekToTime(sinkData: SinkData, time: number): Promise { + try { + if (sinkData.prefetching && sinkData.prefetchPromise) await sinkData.prefetchPromise + if (sinkData.iterator) { + await sinkData.iterator.return() + sinkData.iterator = null + } + sinkData.nextFrame = null + sinkData.iterator = sinkData.sink.canvases(time) + sinkData.lastTime = time + const { value: frame } = await sinkData.iterator.next() + if (frame) { + sinkData.currentFrame = frame + return frame + } + } catch (error) { + console.warn("[editai] failed to seek video:", error) + } + return null + } + + private startPrefetch(sinkData: SinkData) { + if (sinkData.prefetching || !sinkData.iterator || sinkData.nextFrame) return + sinkData.prefetching = true + sinkData.prefetchPromise = this.prefetchNextFrame(sinkData) + } + + private async prefetchNextFrame(sinkData: SinkData): Promise { + if (!sinkData.iterator) { + sinkData.prefetching = false + sinkData.prefetchPromise = null + return + } + try { + const { value: frame, done } = await sinkData.iterator.next() + if (!done && frame) sinkData.nextFrame = frame + } catch (error) { + console.warn("[editai] prefetch failed:", error) + sinkData.iterator = null + } finally { + sinkData.prefetching = false + sinkData.prefetchPromise = null + } + } + + private ensureSink(name: string): Promise { + const existing = this.initPromises.get(name) + if (existing) return existing + const promise = (async () => { + const track = await inputFor(name).getPrimaryVideoTrack() + if (!track) throw new Error(`${name} has no video track.`) + if (!(await track.canDecode())) throw new Error(`This browser cannot decode ${name} (${track.codec ?? "unknown codec"}).`) + this.sinks.set(name, { + sink: new CanvasSink(track, { poolSize: 2 }), + iterator: null, + currentFrame: null, + nextFrame: null, + lastTime: 0, + prefetching: false, + prefetchPromise: null, + }) + })() + // A failed init must not be cached, or the media can never recover. + this.initPromises.set( + name, + promise.catch((err) => { + this.initPromises.delete(name) + throw err + }), + ) + return this.initPromises.get(name)! + } +} + +const isFrameValid = (frame: WrappedCanvas, time: number) => time >= frame.timestamp && time < frame.timestamp + frame.duration + +export const videoCache = new VideoCache() diff --git a/apps/web/src/lib/agent.ts b/apps/web/src/lib/agent.ts new file mode 100644 index 0000000..42f0aa0 --- /dev/null +++ b/apps/web/src/lib/agent.ts @@ -0,0 +1,14 @@ +/** Where the EditAI agent server lives. It owns the project state and the media on disk. */ +export const AGENT_URL = (import.meta.env.VITE_EDITAI_AGENT_URL as string | undefined) ?? "http://localhost:8941" + +/** Media is served with range support, so decoders can seek without downloading the whole file. */ +export const mediaUrl = (name: string) => `${AGENT_URL}/media/${encodeURIComponent(name)}` + +export async function agentJson(path: string, init?: RequestInit): Promise { + const res = await fetch(`${AGENT_URL}${path}`, init) + if (!res.ok) { + const body = await res.text().catch(() => "") + throw new Error(`${init?.method ?? "GET"} ${path} failed with ${res.status}${body ? `: ${body.slice(0, 200)}` : ""}`) + } + return (await res.json()) as T +} diff --git a/apps/web/src/lib/async.ts b/apps/web/src/lib/async.ts new file mode 100644 index 0000000..ab0610b --- /dev/null +++ b/apps/web/src/lib/async.ts @@ -0,0 +1,12 @@ +/** + * Start work nobody is waiting for, without letting a rejection escape unhandled. + * + * The `void` operator would do the same thing more tersely and is banned by the repo's + * conventions, but it also drops rejections on the floor, which this does not. + */ +export function fireAndForget(promise: Promise | undefined, onError?: (error: unknown) => void): void { + promise?.catch((error: unknown) => { + if (onError) onError(error) + else console.warn("[editai] background task failed:", error) + }) +} diff --git a/apps/web/src/routes/index.tsx b/apps/web/src/routes/index.tsx index f086b65..25167ad 100644 --- a/apps/web/src/routes/index.tsx +++ b/apps/web/src/routes/index.tsx @@ -1,5 +1,5 @@ import { createFileRoute } from "@tanstack/react-router" -import { useEffect, useState } from "react" +import { useCallback, useEffect, useState } from "react" import { CommandBar } from "#/components/editor/command-bar" import { Preview } from "#/components/editor/preview" import { RightRail } from "#/components/editor/right-rail" @@ -7,8 +7,12 @@ import { SidePanel } from "#/components/editor/side-panel" import { Timeline } from "#/components/editor/timeline" import { TopBar } from "#/components/editor/top-bar" import { useAssistant } from "#/components/editor/use-assistant" +import { useMediaImport } from "#/components/editor/use-media-import" import { usePlayback } from "#/components/editor/use-playback" import { useProject } from "#/components/editor/use-project" +import { useRenderWorker } from "#/components/editor/use-render-worker" +import { agentJson } from "#/lib/agent" +import { fireAndForget } from "#/lib/async" export const Route = createFileRoute("/")({ component: Editor }) @@ -18,8 +22,29 @@ function Editor() { const { time, playing, toggle, seek, nudge } = usePlayback(project.duration) const [selectedId, setSelectedId] = useState(null) const [prompt, setPrompt] = useState("") + const { imports, importFiles } = useMediaImport() + const { state: render, error: renderError } = useRenderWorker(project) + const [dragging, setDragging] = useState(false) const selected = project.clips.find((c) => c.id === selectedId) ?? null + const media = project.media ?? {} + const exports = project.exports ?? [] + const lastExport = [...exports].reverse().find((e) => e.status === "done") + // Nothing renders without media on disk, so the button says so rather than failing later. + const canExport = + connected && + project.clips.length > 0 && + project.clips.every((c) => c.kind === "text" || Boolean(media[c.name]?.file)) + + const queueExport = useCallback(() => { + fireAndForget( + agentJson("/exports", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ format: "mp4", resolution: "1080p" }), + }), + ) + }, []) // A clip the agent deleted should not stay selected. useEffect(() => { @@ -50,11 +75,38 @@ function Editor() { }, [toggle, seek, nudge, project.duration, project.fps]) return ( -
- +
{ + if (!e.dataTransfer.types.includes("Files")) return + e.preventDefault() + setDragging(true) + }} + onDragLeave={(e) => { + if (e.currentTarget.contains(e.relatedTarget as Node | null)) return + setDragging(false) + }} + onDrop={(e) => { + if (!e.dataTransfer.types.includes("Files")) return + e.preventDefault() + setDragging(false) + fireAndForget(importFiles(Array.from(e.dataTransfer.files))) + }} + > +
- + fireAndForget(importFiles(files))} onSuggest={setPrompt} />
@@ -76,6 +128,12 @@ function Editor() { onSeek={seek} />
+ + {dragging && ( +
+

Drop video or audio to import

+
+ )}
) } diff --git a/bun.lock b/bun.lock index 4a67f7b..ec85304 100644 --- a/bun.lock +++ b/bun.lock @@ -47,6 +47,7 @@ "embla-carousel-react": "^8.6.0", "input-otp": "^1.4.2", "lucide-react": "^1.14.0", + "mediabunny": "^1.55.2", "next-themes": "^0.4.6", "radix-ui": "^1.4.3", "react": "^19.2.0", @@ -79,6 +80,14 @@ "wrangler": "^4.70.0", }, }, + "packages/ffmpeg-sandbox": { + "name": "@editai/ffmpeg-sandbox", + "dependencies": { + "@modelcontextprotocol/sdk": "^1.12.0", + "express": "^4.21.2", + "zod": "^3.24.1", + }, + }, }, "packages": { "@acemir/cssom": ["@acemir/cssom@0.9.31", "", {}, "sha512-ZnR3GSaH+/vJ0YlHau21FjfLYjMpYVIzTD8M8vIEQvIGxeOXyXdzCI140rrCY862p/C/BbzWsjc1dgnM9mkoTA=="], @@ -193,6 +202,8 @@ "@editai/agent": ["@editai/agent@workspace:apps/agent"], + "@editai/ffmpeg-sandbox": ["@editai/ffmpeg-sandbox@workspace:packages/ffmpeg-sandbox"], + "@editai/web": ["@editai/web@workspace:apps/web"], "@emnapi/core": ["@emnapi/core@1.10.0", "", { "dependencies": { "@emnapi/wasi-threads": "1.2.1", "tslib": "^2.4.0" } }, "sha512-yq6OkJ4p82CAfPl0u9mQebQHKPJkY7WrIuk205cTYnYe+k2Z8YBh11FrbRG/H6ihirqcacOgl2BIO8oyMQLeXw=="], @@ -613,23 +624,23 @@ "@tanstack/react-query": ["@tanstack/react-query@5.100.9", "", { "dependencies": { "@tanstack/query-core": "5.100.9" }, "peerDependencies": { "react": "^18 || ^19" } }, "sha512-Oa44XkaI3kCNN6ME0KByU3xT3SEUNOMfZpHxL6+wFoTm+OeUFYHKdeYVe0aOXlRDm/f15sgLwEt2HDorIdW8+A=="], - "@tanstack/react-router": ["@tanstack/react-router@1.170.31", "", { "dependencies": { "@tanstack/history": "1.162.1", "@tanstack/react-store": "^0.9.3", "@tanstack/router-core": "1.171.26", "isbot": "^5.1.22" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-jqITLcf9Y+es9Wm7fD7XfXFGKk1Fujm+okGqHtr+UhISnQuyLQ8d3frsJlIN0Np1doYoeb5weozSkv72+EZ5bw=="], + "@tanstack/react-router": ["@tanstack/react-router@1.170.32", "", { "dependencies": { "@tanstack/history": "1.162.1", "@tanstack/react-store": "^0.9.3", "@tanstack/router-core": "1.171.27", "isbot": "^5.1.22" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-SIpxvaTKco100a5ZR3ePmArbhtm3XOx+w1dpGYY9gxHDta4iXSKDdQuhLonwJbIMkVJsU1rwXf0UDHMrF/1snw=="], "@tanstack/react-router-devtools": ["@tanstack/react-router-devtools@1.167.1", "", { "dependencies": { "@tanstack/router-devtools-core": "1.168.1" }, "peerDependencies": { "@tanstack/react-router": "^1.170.19", "@tanstack/router-core": "^1.171.16", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" }, "optionalPeers": ["@tanstack/router-core"] }, "sha512-pjfGrmjj4d7naEPM7oshqFfwBoxDPNo/UxltlHH5ePbHsJ+plBhd+JaAewm1ueYOjZ0js9hckjWWDYXpCrSfKw=="], "@tanstack/react-router-ssr-query": ["@tanstack/react-router-ssr-query@1.167.1", "", { "dependencies": { "@tanstack/router-ssr-query-core": "1.169.1" }, "peerDependencies": { "@tanstack/query-core": ">=5.90.0", "@tanstack/react-query": ">=5.90.0", "@tanstack/react-router": ">=1.127.0", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-W9j5JPnBikyafvuUfykFfHIWod58OAbAAa5leNkXBcoDoocghMmu6w9uZOmUZvAWT7CSvgj5tBUtF7CM2OoHXQ=="], - "@tanstack/react-start": ["@tanstack/react-start@1.168.48", "", { "dependencies": { "@tanstack/react-router": "1.170.31", "@tanstack/react-start-client": "1.168.29", "@tanstack/react-start-rsc": "0.1.47", "@tanstack/react-start-server": "1.167.36", "@tanstack/router-utils": "1.162.2", "@tanstack/start-client-core": "1.170.26", "@tanstack/start-plugin-core": "1.171.38", "@tanstack/start-server-core": "1.169.30", "pathe": "^2.0.3" }, "peerDependencies": { "@rsbuild/core": "^2.0.0", "@vitejs/plugin-rsc": "*", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0", "vite": ">=7.0.0" }, "optionalPeers": ["@rsbuild/core", "@vitejs/plugin-rsc", "vite"] }, "sha512-ANBtq/phgjPzfwgnB40kyREQGhrLONueJg9y+oFuGIuUVaTfqWjX5y820jsJ0FiZ94JU+WIV6GebVhphNhuYjg=="], + "@tanstack/react-start": ["@tanstack/react-start@1.168.49", "", { "dependencies": { "@tanstack/react-router": "1.170.32", "@tanstack/react-start-client": "1.168.30", "@tanstack/react-start-rsc": "0.1.48", "@tanstack/react-start-server": "1.167.37", "@tanstack/router-utils": "1.162.2", "@tanstack/start-client-core": "1.170.27", "@tanstack/start-plugin-core": "1.171.39", "@tanstack/start-server-core": "1.169.31", "pathe": "^2.0.3" }, "peerDependencies": { "@rsbuild/core": "^2.0.0", "@vitejs/plugin-rsc": "*", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0", "vite": ">=7.0.0" }, "optionalPeers": ["@rsbuild/core", "@vitejs/plugin-rsc", "vite"] }, "sha512-iQb1ZoEHqvZMGLR4G7v3tTtgL06yv/bwxvGP3waVHxVn7bRpyopM44YbOluaGkcJzc13ZTvdLKWYNAN3bxMX/Q=="], - "@tanstack/react-start-client": ["@tanstack/react-start-client@1.168.29", "", { "dependencies": { "@tanstack/react-router": "1.170.31", "@tanstack/router-core": "1.171.26", "@tanstack/start-client-core": "1.170.26" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-YhEtx3EwbsLUWXSdwclz/C/AlDqPNoNMoNTEHulYDruaMAKNbFeWsDgBbHv/aE6P2P3shAMFYAT1cVHZM1WcJQ=="], + "@tanstack/react-start-client": ["@tanstack/react-start-client@1.168.30", "", { "dependencies": { "@tanstack/react-router": "1.170.32", "@tanstack/router-core": "1.171.27", "@tanstack/start-client-core": "1.170.27" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-qsZuykUl1EF0/rc1bin1RtjFzz07YMOTBzhOstSDzbOVm/WKf1QKFTN+qAZi74xlbXSWNuT2MiFRyg4RKwj8iw=="], - "@tanstack/react-start-rsc": ["@tanstack/react-start-rsc@0.1.47", "", { "dependencies": { "@tanstack/react-router": "1.170.31", "@tanstack/router-core": "1.171.26", "@tanstack/router-utils": "1.162.2", "@tanstack/start-client-core": "1.170.26", "@tanstack/start-fn-stubs": "1.162.0", "@tanstack/start-plugin-core": "1.171.38", "@tanstack/start-storage-context": "1.167.28", "pathe": "^2.0.3" }, "peerDependencies": { "@rspack/core": ">=2.0.0-0", "@vitejs/plugin-rsc": ">=0.5.30", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0", "react-server-dom-rspack": ">=0.0.2" }, "optionalPeers": ["@rspack/core", "@vitejs/plugin-rsc", "react-server-dom-rspack"] }, "sha512-Tr01FbuRRpHSwkLjyfg7OotXgrJ4ufDjGKuk9HpLe7a2ouEs1Iq/kC2RKY0alvCYX+cpYFVayoyxNZCeybCTNQ=="], + "@tanstack/react-start-rsc": ["@tanstack/react-start-rsc@0.1.48", "", { "dependencies": { "@tanstack/react-router": "1.170.32", "@tanstack/router-core": "1.171.27", "@tanstack/router-utils": "1.162.2", "@tanstack/start-client-core": "1.170.27", "@tanstack/start-fn-stubs": "1.162.0", "@tanstack/start-plugin-core": "1.171.39", "@tanstack/start-storage-context": "1.167.29", "pathe": "^2.0.3" }, "peerDependencies": { "@rspack/core": ">=2.0.0-0", "@vitejs/plugin-rsc": ">=0.5.30", "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0", "react-server-dom-rspack": ">=0.0.2" }, "optionalPeers": ["@rspack/core", "@vitejs/plugin-rsc", "react-server-dom-rspack"] }, "sha512-UglRdTMuF3c4dvzL/gh4dMVbMWHsPy8ZgTQdT2qpTlyt1b/3m+R40tBFdsfTgw76VTXivW+qV/ih0PN9000XTw=="], - "@tanstack/react-start-server": ["@tanstack/react-start-server@1.167.36", "", { "dependencies": { "@tanstack/react-router": "1.170.31", "@tanstack/router-core": "1.171.26", "@tanstack/start-server-core": "1.169.30" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-v22xUhNENr0u67AIDbfbJgmQwMGqcrwdqqNI44Nd605G4DGTA7W1KPbea8bBjPkHxf5I+FTKL8Nb56w2Sb/Gpw=="], + "@tanstack/react-start-server": ["@tanstack/react-start-server@1.167.37", "", { "dependencies": { "@tanstack/react-router": "1.170.32", "@tanstack/router-core": "1.171.27", "@tanstack/start-server-core": "1.169.31" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-cODHpFU8vIm7AdHii9W3NEuwyruNmT5wLDZjRsJLT9Jp+7FACnfJrRbvxnp1ldSv/9mxcHKi/OgwbdHQeMZcAQ=="], "@tanstack/react-store": ["@tanstack/react-store@0.9.3", "", { "dependencies": { "@tanstack/store": "0.9.3", "use-sync-external-store": "^1.6.0" }, "peerDependencies": { "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0", "react-dom": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0" } }, "sha512-y2iHd/N9OkoQbFJLUX1T9vbc2O9tjH0pQRgTcx1/Nz4IlwLvkgpuglXUx+mXt0g5ZDFrEeDnONPqkbfxXJKwRg=="], - "@tanstack/router-core": ["@tanstack/router-core@1.171.26", "", { "dependencies": { "@tanstack/history": "1.162.1", "cookie-es": "^3.0.0", "seroval": "^1.6.2", "seroval-plugins": "^1.6.2" } }, "sha512-VymmPSs/93szHur/7PBygT6tRbmfcNO1xR46Wn/JVHeBqVduLPCuxBeJEgfVmVq8E4hSklPEh8i6qGXQd6WRqA=="], + "@tanstack/router-core": ["@tanstack/router-core@1.171.27", "", { "dependencies": { "@tanstack/history": "1.162.1", "cookie-es": "^3.0.0", "seroval": "^1.6.2", "seroval-plugins": "^1.6.2" } }, "sha512-wDwSLvoLwIaNcnx9UNcN9Mb7Y8QwCYq1U1RQZwyN186gnkIoIYI2SOxy8VqH1vFigbkHkk4FmwMAQlghPgDK2g=="], "@tanstack/router-devtools-core": ["@tanstack/router-devtools-core@1.168.1", "", { "dependencies": { "clsx": "^2.1.1", "goober": "^2.1.16" }, "peerDependencies": { "@tanstack/router-core": "^1.171.16", "csstype": "^3.0.10" }, "optionalPeers": ["csstype"] }, "sha512-qr4voa4cpSMwQvS3867xkU3AB3MtJbTuovKIy+btjJ/Faju6er9w0nDylmD+005Mk/3YKw9/iueZJl2JAB7JOA=="], @@ -641,15 +652,15 @@ "@tanstack/router-utils": ["@tanstack/router-utils@1.162.2", "", { "dependencies": { "@babel/generator": "^7.28.5", "@babel/parser": "^7.28.5", "@babel/types": "^7.28.5", "ansis": "^4.1.0", "babel-dead-code-elimination": "^1.0.12", "diff": "^8.0.2", "pathe": "^2.0.3", "tinyglobby": "^0.2.15" } }, "sha512-hTWqJtqIFFdvuCl8WXNyrodp2L9zo2G37xKRrcVmVRWpAB2h+U1LuRAfS4tsFTiWOIoE/B+WDVFB8JpoEdw6jQ=="], - "@tanstack/start-client-core": ["@tanstack/start-client-core@1.170.26", "", { "dependencies": { "@tanstack/router-core": "1.171.26", "@tanstack/start-fn-stubs": "1.162.0", "@tanstack/start-storage-context": "1.167.28", "seroval": "^1.6.2" } }, "sha512-a8SmEQ4gIeWDHBSqy1ePFNwXu0kHTPCoMiEdE0eKRBITVsexfEX8yKvqxLi66cF6aGCmDIkndbth3yB7R344zQ=="], + "@tanstack/start-client-core": ["@tanstack/start-client-core@1.170.27", "", { "dependencies": { "@tanstack/router-core": "1.171.27", "@tanstack/start-fn-stubs": "1.162.0", "@tanstack/start-storage-context": "1.167.29", "seroval": "^1.6.2" } }, "sha512-Ro6ZSM0NgYKDMxM0e8qyU4mBfnld2Zb74BA/9f4i35C0Y3IAA8Zxs/DIPfOo9VRYWp5c5n5Q8dRBn2qn0vtPhQ=="], "@tanstack/start-fn-stubs": ["@tanstack/start-fn-stubs@1.162.0", "", {}, "sha512-QWfUZ3Yo923tdQn38LyKMU8rcTw69zc+T4dAvgTWV4O56SqFRsGfS0lSWIMhJRwXIx/bvdi7nTUBDdZtTHtpTQ=="], - "@tanstack/start-plugin-core": ["@tanstack/start-plugin-core@1.171.38", "", { "dependencies": { "@babel/code-frame": "7.27.1", "@babel/core": "^7.28.5", "@babel/types": "^7.28.5", "@tanstack/router-core": "1.171.26", "@tanstack/router-generator": "1.167.32", "@tanstack/router-plugin": "1.168.34", "@tanstack/router-utils": "1.162.2", "@tanstack/start-server-core": "1.169.30", "exsolve": "^1.0.7", "lightningcss": "^1.32.0", "pathe": "^2.0.3", "picomatch": "^4.0.3", "seroval": "^1.6.2", "source-map": "^0.7.6", "srvx": "^0.11.9", "tinyglobby": "^0.2.15", "ufo": "^1.5.4", "vitefu": "^1.1.1", "xmlbuilder2": "^4.0.3", "zod": "^4.4.3" }, "peerDependencies": { "@rsbuild/core": "^2.0.0", "vite": ">=7.0.0" }, "optionalPeers": ["@rsbuild/core", "vite"] }, "sha512-ldRjxosoTyWG0xP6CLo/DkiuENt668se7oZwAJuRMtRBGBDzhMVMwFnS86P2qgnpj/BQVFebVpXMVfzeH7tvDw=="], + "@tanstack/start-plugin-core": ["@tanstack/start-plugin-core@1.171.39", "", { "dependencies": { "@babel/code-frame": "7.27.1", "@babel/core": "^7.28.5", "@babel/types": "^7.28.5", "@tanstack/router-core": "1.171.27", "@tanstack/router-generator": "1.167.33", "@tanstack/router-plugin": "1.168.35", "@tanstack/router-utils": "1.162.2", "@tanstack/start-server-core": "1.169.31", "exsolve": "^1.0.7", "lightningcss": "^1.32.0", "pathe": "^2.0.3", "picomatch": "^4.0.3", "seroval": "^1.6.2", "source-map": "^0.7.6", "srvx": "^0.11.9", "tinyglobby": "^0.2.15", "ufo": "^1.5.4", "vitefu": "^1.1.1", "xmlbuilder2": "^4.0.3", "zod": "^4.4.3" }, "peerDependencies": { "@rsbuild/core": "^2.0.0", "vite": ">=7.0.0" }, "optionalPeers": ["@rsbuild/core", "vite"] }, "sha512-Zyj6G4MDFLXcHYhPavhewCBo8dsxi3qvPk31zl/QtTzYzOY9rJHDDBGarKsJq20ziChmkGn/sttYqb4vGCxJDA=="], - "@tanstack/start-server-core": ["@tanstack/start-server-core@1.169.30", "", { "dependencies": { "@tanstack/history": "1.162.1", "@tanstack/router-core": "1.171.26", "@tanstack/start-client-core": "1.170.26", "@tanstack/start-storage-context": "1.167.28", "fetchdts": "^0.1.6", "h3-v2": "npm:h3@2.0.1-rc.20", "seroval": "^1.6.2" } }, "sha512-knlKOPsHyDbGKgP1pfUowtObZIKS9zRI6vl7dzbA8D/NJo6dMdDdqGcDf4PINqREW8hCHdkB4At0y1/Mk+L2SQ=="], + "@tanstack/start-server-core": ["@tanstack/start-server-core@1.169.31", "", { "dependencies": { "@tanstack/history": "1.162.1", "@tanstack/router-core": "1.171.27", "@tanstack/start-client-core": "1.170.27", "@tanstack/start-storage-context": "1.167.29", "fetchdts": "^0.1.6", "h3-v2": "npm:h3@2.0.1-rc.20", "seroval": "^1.6.2" } }, "sha512-56w8l+Fao01YCrmv0hzNxL3b3FRmqLRGdE11izZNQsW+1CcSYuehAM5khWKABuI1DI/UIBLGV7BwL4Dlg0eHCw=="], - "@tanstack/start-storage-context": ["@tanstack/start-storage-context@1.167.28", "", { "dependencies": { "@tanstack/router-core": "1.171.26" } }, "sha512-ZGl6bEXW5m77VVtWMt/naEBVoIeQkXYgF4UgiJM/3z3e1wbRzaoW7Dnc4HW+hoBPOOdfiVVNx31RT2ARnXWAiQ=="], + "@tanstack/start-storage-context": ["@tanstack/start-storage-context@1.167.29", "", { "dependencies": { "@tanstack/router-core": "1.171.27" } }, "sha512-8qfprC5774XMRDQlMogkfiGpFLiBf0xDG4bMFUbfkSzpCAQwpLbgGY4Zwft22O9rKYq2vUXusKAdVjvJTUEquQ=="], "@tanstack/store": ["@tanstack/store@0.9.3", "", {}, "sha512-8reSzl/qGWGGVKhBoxXPMWzATSbZLZFWhwBAFO9NAyp0TxzfBP0mIrGb8CP8KrQTmvzXlR/vFPPUrHTLBGyFyw=="], @@ -707,6 +718,10 @@ "@types/deep-eql": ["@types/deep-eql@4.0.2", "", {}, "sha512-c9h9dVVMigMPc4bwTvC5dxqtqJZwQPePsWjPlpSOnojbor6pGqdk541lfA7AqFQr5pB1BRdq0juY9db81BwyFw=="], + "@types/dom-mediacapture-transform": ["@types/dom-mediacapture-transform@0.1.12", "", { "dependencies": { "@types/dom-webcodecs": "*" } }, "sha512-d7/QsLRwF864A5mgIM/YrfiglHoYn7zgCcAoJgW404r+2DwnNr7EBbLnCWpmOMgH8y0te73L1AV6H1bmauaWFw=="], + + "@types/dom-webcodecs": ["@types/dom-webcodecs@0.1.13", "", {}, "sha512-O5hkiFIcjjszPIYyUSyvScyvrBoV3NOEEZx/pMlsu44TKzWNkLVBBxnxJz42in5n3QIolYOcBYFCPZZ0h8SkwQ=="], + "@types/estree": ["@types/estree@1.0.9", "", {}, "sha512-GhdPgy1el4/ImP05X05Uw4cw2/M93BCUmnEvWZNStlCzEKME4Fkk+YpoA5OiHNQmoS7Cafb8Xa3Pya8m1Qrzeg=="], "@types/express": ["@types/express@4.17.25", "", { "dependencies": { "@types/body-parser": "*", "@types/express-serve-static-core": "^4.17.33", "@types/qs": "*", "@types/serve-static": "^1" } }, "sha512-dVd04UKsfpINUnK0yBoYHDF3xu7xVH4BuDotC/xGuycx4CgbP48X/KF/586bcObxT0HENHXEU8Nqtu6NR+eKhw=="], @@ -1207,6 +1222,8 @@ "media-typer": ["media-typer@0.3.0", "", {}, "sha512-dq+qelQ9akHpcOl/gUVRTxVIOkAJ1wR3QAvb4RsVjS8oVoFjDGTc679wJYmUmknUF5HwMLOgb5O+a3KxfWapPQ=="], + "mediabunny": ["mediabunny@1.55.2", "", { "dependencies": { "@types/dom-mediacapture-transform": "^0.1.11", "@types/dom-webcodecs": "0.1.13" } }, "sha512-EEx4O6qYddAdCyWPMZNDwI7uc5hewNHrPAf9jLcVhIbXoPsiqNQ+D9i1pfadmGkjN2V318jSrZljkpoziYm6Lg=="], + "merge-descriptors": ["merge-descriptors@1.0.3", "", {}, "sha512-gaNvAS7TZ897/rVaZ0nMtAyxNyi/pdbjbAwUpFQpN70GqnVfOiXpeUUMKRBmzXaSQ8DdTX4/0ms62r2K+hE6mQ=="], "merge-stream": ["merge-stream@2.0.0", "", {}, "sha512-abv/qOcuPfk3URPfDzmZU1LKmuw8kT+0nIHvKrKgFrwifol/doWcdA4ZqsWQ8ENrFKkd67Mfpo/LovbIUsbt3w=="], @@ -1643,12 +1660,20 @@ "@tanstack/devtools-event-bus/ws": ["ws@8.20.0", "", { "peerDependencies": { "bufferutil": "^4.0.1", "utf-8-validate": ">=5.0.2" }, "optionalPeers": ["bufferutil", "utf-8-validate"] }, "sha512-sAt8BhgNbzCtgGbt2OxmpuryO63ZoDk/sqaB/znQm94T4fCEsy/yV+7CdC1kJhOU9lboAEU7R3kquuycDoibVA=="], + "@tanstack/router-generator/@tanstack/router-core": ["@tanstack/router-core@1.171.26", "", { "dependencies": { "@tanstack/history": "1.162.1", "cookie-es": "^3.0.0", "seroval": "^1.6.2", "seroval-plugins": "^1.6.2" } }, "sha512-VymmPSs/93szHur/7PBygT6tRbmfcNO1xR46Wn/JVHeBqVduLPCuxBeJEgfVmVq8E4hSklPEh8i6qGXQd6WRqA=="], + "@tanstack/router-generator/zod": ["zod@4.4.3", "", {}, "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ=="], + "@tanstack/router-plugin/@tanstack/router-core": ["@tanstack/router-core@1.171.26", "", { "dependencies": { "@tanstack/history": "1.162.1", "cookie-es": "^3.0.0", "seroval": "^1.6.2", "seroval-plugins": "^1.6.2" } }, "sha512-VymmPSs/93szHur/7PBygT6tRbmfcNO1xR46Wn/JVHeBqVduLPCuxBeJEgfVmVq8E4hSklPEh8i6qGXQd6WRqA=="], + "@tanstack/router-plugin/zod": ["zod@4.4.3", "", {}, "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ=="], "@tanstack/start-plugin-core/@babel/code-frame": ["@babel/code-frame@7.27.1", "", { "dependencies": { "@babel/helper-validator-identifier": "^7.27.1", "js-tokens": "^4.0.0", "picocolors": "^1.1.1" } }, "sha512-cjQ7ZlQ0Mv3b47hABuTevyTuYN4i+loJKGeV9flcCgIK37cCXRh+L1bd3iBHlynerhQ7BhCkn2BPbQUL+rGqFg=="], + "@tanstack/start-plugin-core/@tanstack/router-generator": ["@tanstack/router-generator@1.167.33", "", { "dependencies": { "@babel/types": "^7.28.5", "@tanstack/router-core": "1.171.27", "@tanstack/router-utils": "1.162.2", "@tanstack/virtual-file-routes": "1.162.0", "jiti": "^2.7.0", "magic-string": "^0.30.21", "prettier": "^3.5.0", "zod": "^4.4.3" } }, "sha512-Z3lCWIPuRUMPmuI8Mm48x/s49TxmHOaFVZ52j1W1QKYrsFHyT6U/h9bqfHJDxfQ8kz7y9q+W1YPKZ15Ee7yuCA=="], + + "@tanstack/start-plugin-core/@tanstack/router-plugin": ["@tanstack/router-plugin@1.168.35", "", { "dependencies": { "@babel/core": "^7.28.5", "@babel/template": "^7.27.2", "@babel/types": "^7.28.5", "@tanstack/router-core": "1.171.27", "@tanstack/router-generator": "1.167.33", "@tanstack/router-utils": "1.162.2", "chokidar": "^5.0.0", "unplugin": "^3.0.0", "zod": "^4.4.3" }, "peerDependencies": { "@rsbuild/core": ">=1.0.2 || ^2.0.0", "@tanstack/react-router": "^1.170.32", "vite": ">=5.0.0 || >=6.0.0 || >=7.0.0 || >=8.0.0", "vite-plugin-solid": "^2.11.10 || ^3.0.0-0", "webpack": ">=5.92.0" }, "optionalPeers": ["@rsbuild/core", "@tanstack/react-router", "vite", "vite-plugin-solid", "webpack"] }, "sha512-foDAZKFqHXae+oFbIgcsSvy2QCVRn7XdS3nhwcRvD+ed6JrKPUP/1lQMsZLJqWycgR1vkZF7gs955KGa0NZQ0w=="], + "@tanstack/start-plugin-core/zod": ["zod@4.4.3", "", {}, "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ=="], "accepts/mime-types": ["mime-types@2.1.35", "", { "dependencies": { "mime-db": "1.52.0" } }, "sha512-ZDY+bPm5zTTF+YpCrAU9nK0UgICYPT0QtT1NZWFv4s++TNkcgVaT0g6+4R2uI4MjQjzysHB1zxuWL50hzaeXiw=="], diff --git a/package.json b/package.json index 76881f5..81afb25 100644 --- a/package.json +++ b/package.json @@ -1,7 +1,7 @@ { "name": "editai", "private": true, - "workspaces": ["apps/*"], + "workspaces": ["apps/*", "packages/*"], "scripts": { "dev": "turbo dev", "dev:web": "turbo dev --filter=@editai/web", diff --git a/packages/ffmpeg-sandbox/mcp-ffmpeg-sandbox.mjs b/packages/ffmpeg-sandbox/mcp-ffmpeg-sandbox.mjs index 8dd5a49..038829e 100644 --- a/packages/ffmpeg-sandbox/mcp-ffmpeg-sandbox.mjs +++ b/packages/ffmpeg-sandbox/mcp-ffmpeg-sandbox.mjs @@ -4,12 +4,12 @@ // Every ffmpeg/ffprobe invocation runs inside a throwaway Docker container with // no network, a single mounted workspace, capped memory/CPU, and a hard timeout. // -// npm i @modelcontextprotocol/sdk express zod -// docker build -t editai-sandbox . -// node mcp-ffmpeg-sandbox.mjs +// bun install (from the repo root; this is a workspace package) +// bun run build:image (docker build -f Dockerfile.sandbox -t editai-sandbox .) +// bun run start (http://localhost:8931/mcp) // -// Then in TrueForge: Settings -> Connectors -> add server by URL -// http://localhost:8931/mcp +// `bun run setup` in apps/agent finds it here via /health, registers it with +// TrueForge as the "ffmpeg-sandbox" connector, and attaches it to the agent. import express from "express"; import { z } from "zod"; diff --git a/packages/ffmpeg-sandbox/package.json b/packages/ffmpeg-sandbox/package.json new file mode 100644 index 0000000..39185c0 --- /dev/null +++ b/packages/ffmpeg-sandbox/package.json @@ -0,0 +1,16 @@ +{ + "name": "@editai/ffmpeg-sandbox", + "private": true, + "type": "module", + "description": "MCP server exposing a dockerized media workbench: every ffmpeg/ffprobe/python invocation runs in a throwaway container with no network.", + "scripts": { + "build:image": "docker build -f Dockerfile.sandbox -t editai-sandbox .", + "start": "node mcp-ffmpeg-sandbox.mjs", + "dev": "node --watch mcp-ffmpeg-sandbox.mjs" + }, + "dependencies": { + "@modelcontextprotocol/sdk": "^1.12.0", + "express": "^4.21.2", + "zod": "^3.24.1" + } +} diff --git a/site/.gitignore b/site/.gitignore new file mode 100644 index 0000000..e985853 --- /dev/null +++ b/site/.gitignore @@ -0,0 +1 @@ +.vercel diff --git a/site/README.md b/site/README.md new file mode 100644 index 0000000..e4d7e00 --- /dev/null +++ b/site/README.md @@ -0,0 +1,37 @@ +# site + +The EditAI landing page: + +One hand-written `index.html` with inline CSS and JS. No build step, no dependencies, no +framework. Open the file to work on it. + +```bash +python3 -m http.server 4477 --directory site # http://localhost:4477 +``` + +## The hero timeline + +The timeline in the hero is not a screenshot. It renders the same sample project the agent server +ships with (`apps/agent/src/project.ts`: `intro.mp4`, `b-roll.mp4`, `talking-head.mp4`, the +voiceover and its three silences) and performs a real ripple delete on it: `mapTime()` in the page +script applies the same rule the server does, so 24.0s becomes 21.1s and clips straddling a silence +get shorter rather than merely shifting. + +**If the sample project changes, change `SILENCES` and `LANES` in `index.html` to match.** The page +claims those are real numbers, so they have to stay real. + +## Colours + +Every colour is lifted from the editor rather than invented: the well, panel and foreground greys +from `apps/web/src/styles.css`, the violet `#7c5cff` from the favicon, and the three clip colours +(`#3aa39b` video, `#e0a63b` text, `#5fae63` audio) that the timeline paints tracks with. + +## Deploying + +```bash +cd site +vercel deploy --prod --scope deonmenezes-projects +``` + +Production aliases: `editai-agent.vercel.app` (canonical), `edit-ai-video.vercel.app`, +`edit-ai-lemon.vercel.app`. diff --git a/site/index.html b/site/index.html new file mode 100644 index 0000000..a1a5e07 --- /dev/null +++ b/site/index.html @@ -0,0 +1,950 @@ + + + + + +EditAI: talk to your timeline + + + + + + + + + + + + + + + + + + +
+ + + EditAI + + 00:00:00:00 + + + + + GitHub + +
+ + + + +
+
+ + +
+
+ 00:00:00 + Cold open + +
+ +

Talkto yourtimeline

+ +

+ EditAI is an AI harness for video editing. Ask for the cut, the motion + graphic, the filter or the noise gate. An agent makes the edit on your real timeline + through 19 MCP tools, and stops for your approval before anything destructive. +

+ + +
+
+ + + ⌘⏎ +
+ +
+
V1
+
A1
+
T1
+
+ +
+ Waiting for an instruction + + + 24.0s +
+
+ +

+ Real behaviour, real numbers: find_silences then + ripple_delete across every track, 24.0s down to 21.1s, which is + exactly the 2.9 seconds of dead air. Nothing is cut until you approve it, and every edit undoes. +

+
+ + +
+
+ 00:00:14 + The handoff + +
+

Connect the MCP.
The agent takes the room.

+

+ Point your client at EditAI's MCP server and an agent comes up inside a sandbox on + TrueForge, TrueFoundry's open-source agent harness. From there it drives the + edit at whichever layer the job needs. +

+ +
+
+ Layer one +

Your editing UI

+

Bring your own front end. The timeline is exposed as tools, not as a prompt describing a + timeline, so any MCP client can drive it.

+
any MCP client
+
+
+ Layer two +

Our editor

+

A full timeline, preview and assistant panel. The agent's edits stream in over SSE, so + the tracks redraw as the work happens.

+
apps/web · TanStack Start
+
+
+ Layer three +

The encode itself

+

When a job needs the real file, the agent writes the ffmpeg filter graph and runs it in a + locked-down container.

+
ffmpeg · ffprobe · python
+
+
+ +
# 1 · the harness
+npx @truefoundry/trueforge@latest          → localhost:8790
+
+# 2 · the timeline tools
+cd apps/agent && bun run start            → localhost:8941
+
+# 3 · wire them together (any one key is enough)
+ANTHROPIC_API_KEY=sk-... bun run setup
+
+# 4 · the editor
+bun run dev:web                           → localhost:5173
+
+ + +
+
+ 00:00:36 + Say it + +
+

Ask in a sentence.
Watch the tracks move.

+

+ Every ask below runs real tools against real state. Clip ids, source offsets and frame + boundaries are the agent's problem, not yours. +

+ +
+
+

▸Remove the silences

+

Finds every silent range on the voice track, tells you what will go, and + ripple-deletes it across all tracks so the gaps close.

+

find_silences → remove_silences

+
+
+

▸Caption every clip

+

Fans out one sub-agent per clip to transcribe in parallel, merges the + results, and lays timed captions on their own track.

+

transcribe_clip × N → add_captions

+
+
+

▸Cut this to the beat

+

Reads the tempo off the music track and splits on the beat grid, both + halves still frame-accurate.

+

detect_beats → split_clip

+
+
+

▸Duck the music under the voiceover

+

Sets clip volume where the voice track is speaking and puts it back + where it is not.

+

get_project → set_volume

+
+
+

▸Grade it warmer and add grain

+

Writes the ffmpeg filter graph, runs it in the media sandbox, then probes + the output to confirm it is what you asked for.

+

probe_media → run_ffmpeg

+
+
+

▸Kill the room tone

+

Denoises the voice track in the sandbox and leaves the original file + untouched next to it.

+

run_ffmpeg · afftdn

+
+
+ +

+ Sixteen timeline tools in all. Five read, seven write and undo, three destructive, one gated + export. The destructive ones carry MCP's destructiveHint, which is + what makes the harness stop and ask you first. +

+
+ + +
+
+ 00:00:58 + Pull it in + +
+

Need a logo?
It goes and gets one.

+

+ The agent reaches the live web through Bright Data, so a logo, a product shot + or a reference frame is one sentence away from being on your timeline. The connector attaches + deferred: it costs no context until a task actually needs it. +

+
+
+ Find it +

search_engine

+

Live search results, not a training-set memory of what a brand looked like two years ago.

+
+
+ Take it +

scrape_as_markdown

+

Pulls the page down clean, so the agent can lift the asset and the copy around it.

+
+
+
+ + +
+
+ 00:01:16 + Add a tool mid-call + +
+

A new skill is a
line in a config.

+ +
+
+ On the call + Bro, can you add an MCP for motion graphics? I want to use it in this video. +
+
+ You + Sent to your agent. +
+
+ +

+ Every capability here is an MCP server, so adding one is a name in a list and a restart, not a + release. The agent picks up the new tools on its next turn and starts using them in the same + conversation. +

+ +
EDITAI_CONNECTORS=exa,bright-data,motion-graphics bun run setup
+
+✓ motion-graphics   attached   read-only, deferred
+✓ agent updated     14 connectors → tools live on the next turn
+ +

+ Catalog servers attach by name, OAuth ones authorise in chat, and anything you host yourself + attaches the same way. Read-only and deferred by default, so a connector you rarely use does + not crowd the agent's context. +

+
+ + +
+
+ 00:01:34 + Where it runs + +
+

Nothing the agent
writes runs on your box.

+

+ Two sandboxes, because the two kinds of work need different boundaries. +

+
+
+ Daytona +

Harness sandbox

+

For the computation the agent should not estimate: arithmetic, parsing, anything it + would otherwise guess at. Runs remotely, registered the moment the key is present.

+
exit 0 · 338350
+
+
+ Container +

Media sandbox

+

Every ffmpeg call runs in a throwaway container: no network, capped memory and CPU, 256 + pids, non-root, one mounted workspace, hard timeout. Python there asks for approval first, + because a script with the workspace mounted can delete your source media.

+
--network none · --pids-limit 256
+
+
+
+ + +
+
+ 00:01:52 + Reviewed by Qodo + +
+

Every line here was
reviewed by Qodo.

+

+ Not a badge. Eight findings across three reviews, each one either fixed or answered with a + proof, and a merge gate that reads Qodo's structured verdict rather than its prose. +

+ +
+ 8Findings raised + 7Real, fixed + 1Checked, rejected + 5Regression tests added +
+ +
+
+ PR #3 + Duplicate clip ids after repeated ripple deletes. A clip + cut twice produced the same id twice, breaking lookup, deletion and React keys. The test + suite had missed it because the existing test asserted the buggy id as correct. + FIXED +
+
+ PR #3 + Export announced before the file existed. The record was + committed, notifying subscribers, and only then was the file written. + FIXED +
+
+ PR #3 + Session restore leaked its stream on unmount. It set state + on a gone component and left the connection open. + FIXED +
+
+ PR #3 + Setup posted API keys over plaintext HTTP. It now refuses + to do that to anything but localhost. + FIXED +
+
+ PR #3 + The trim bound was called wrong, and was not. The source + expression expands to the correct source time; the code accepts exactly up to the limit + and rejects one frame past it. Two tests now pin that boundary. + REJECTED +
+
+ PR #2 + Sandboxed Python could delete the source media. The + argument guard could not catch it, and the container is no defence, since the workspace is + exactly what it is meant to reach. The tool now asks for approval first. + FIXED +
+
+ +

+ Qodo publishes no check run and no approving review, so GitHub's own auto-merge has nothing to + gate on. The workflow in this repo is that missing gate: it reads only the structured counter + chips, never the prose, because Qodo quotes findings verbatim and a pull request can put the + words "no issues found" into its own diff. It binds the verdict to the reviewed commit and + merges with --match-head-commit, so a push racing the merge is + rejected rather than slipped in. +

+
+ + +
+

And the video you
just watched?

+

It cut itself.

+ +
+ + + +
+
+ + + + diff --git a/site/vercel.json b/site/vercel.json new file mode 100644 index 0000000..92bf592 --- /dev/null +++ b/site/vercel.json @@ -0,0 +1,13 @@ +{ + "$schema": "https://openapi.vercel.sh/vercel.json", + "cleanUrls": true, + "headers": [ + { + "source": "/(.*)", + "headers": [ + { "key": "X-Content-Type-Options", "value": "nosniff" }, + { "key": "Referrer-Policy", "value": "strict-origin-when-cross-origin" } + ] + } + ] +}