Video perception + launch-video creation for Claude Code.
Perception: point it at any video (local file or URL), get a structured analysis the model can reason about: scenes, motion windows, action attribution, transcription, and base64 frames with temporal narrative guidance. Five MCP tools: inspect (cheap forecast), analyze (structural pass), watch (frames + audio), measure (exact token forecast via Anthropic's count_tokens), configure (settings).
Creation (v0.16): scaffold a launch video from a storyboard. A 79-effect canonical library, a studio dashboard (library / composer / preview), a machine-readable design lock, a restage procedure, and three render engines (hyperframes, own, own-parallel) behind one entry point. See Creation.
From the s0nderlabs marketplace:
claude plugin marketplace add s0nderlabs/marketplace
claude plugin install lumiere@s0nderlabsOr local development:
cd ~/Documents/s0nderlabs/lumiere
bun install
c --plugin-dir ~/Documents/s0nderlabs/lumiereffmpeg+ffprobe(required)yt-dlp(URL inputs)whisper-clifromwhisper.cpp(default audio backend) orGEMINI_API_KEYMAX_MCP_OUTPUT_TOKENS=100000recommendedLUMIERE_ANTHROPIC_API_KEY(optional, only for themeasuretool andlumiere-costCLI)- Google Chrome (optional, only for the own render engine:
lumiere-render --engine own)
macOS:
brew install ffmpeg yt-dlp whisper-cpp/lumiere https://x.com/.../status/123
By default the skill fans the watching out to a background workflow (skills/lumiere/perceive.workflow.mjs): a low-tier scout skims the whole video for the narrative arc, then one subagent per segment deep-watches its slice at the configured tier, and the results merge into a structured beat sheet (creation/perception-beatsheet.schema.json). The frame-heavy work stays in the subagents, so your session stays clean, and the beat sheet prefills a creation lock's scenes[]. If the Workflow tool is unavailable it falls back to an inline inspect -> analyze -> chunked watch loop that returns a narrative.
Cheap metadata pass + per-tier context-cost preview. Always call first on a new video. Returns: duration, resolution, codec, fps, audio presence, recommended tier, autocompact warning.
v0.10.3 splits the per-tier estimate into two metrics: mcp_tokens_per_* (chars/3.5 transport metric, predicts per-call MCP truncation) and conversation_tokens_per_* (Anthropic image-token formula, predicts whole-transcript autocompact). The previous combined metric over-warned autocompact at high resolution. Pass exact_tokens=true (requires LUMIERE_ANTHROPIC_API_KEY) to probe one frame per tier and call count_tokens for exact conversation-token values; falls back to the heuristic otherwise.
Structural ffmpeg pass. Returns: scene cuts, silence intervals, motion windows, subject bbox (3-tier cascade: CC segmentation → cropdetect → center-prior heuristic, each with a confidence 0.0-1.0), palette outliers (color novelty), transcription, and (v0.11) content_class (7-way classification: animation | ui-screen | human-motion | talking-head | real-world | nature | generic) plus motion_detection_warning when global motion windows are unreliable (subject-region siti finds no peaks → synthetic-middle fallback fires). No frames extracted. Plan chunks from this.
Frame extraction + audio. Returns base64 images with narrative guidance. Key params:
mode:low(384px) |mid(512px) |high(1024px, default) |max(1536px)narrative_mode: temporal sequence reading instead of frame-by-frameadaptive_sampling: motion-window-aware frame allocation (70/30 motion/static split)roi:"auto"readsanalyze.subject_bbox;"per-window"(v0.8) assigns each motion window its own bbox fromanalyze.window_bboxesso a traveling subject stays tight in every crop (requiresadaptive_sampling=true+ prioranalyzewith motion=true);"x,y,w,h"explicit. Subject gets full target resolutionstart_time/end_time: chunk a sub-segmentview_sample: override the auto-budgetprobe_calibration(v0.10.3, default-on in v0.11): extract one probe frame at the target resolution + crop BEFORE the main pool, measure actual chars/3.5, derive a per-videoview_samplefrom the measurement. Replaces the staticSAFE_AT_100Ktable for that call. Useful for outlier content (dense terminal UI, sparse flat colors). Adds ~150-300ms.LUMIERE_PROBE_CALIBRATION=0disables globally.
watch's audio path is cache-aware (v0.11): if analyze already flagged the transcription as low-confidence (typical for ambient gym / music / quiet audio that whisper hallucinates on), watch skips the whisper re-run and emits the same skipped-reason text. Eliminates the v0.10.x leak where every chunk got fresh whisper hallucinations.
Auto-budget respects MAX_MCP_OUTPUT_TOKENS. Runtime trim drops frames if the response would exceed the cap. The Budget block (containing the gate-verification fields: view_sample_applied, extraction_fps, proactive_sizing, runtime_trim, out_of_range_dropped, etc.) is always emitted, including when skip_metadata=true or narrative_mode=true would suppress the verbose Source/Manifest/Video/Audio dumps.
Exact token forecast via Anthropic's /v1/messages/count_tokens (free endpoint, requires LUMIERE_ANTHROPIC_API_KEY). Extracts frames, builds the would-be response payload, counts tokens for the current Claude model, returns exact conversation_tokens + heuristic mcp_cap_tokens. Discards the payload. Use BEFORE a high-stakes watch call. Auto-detects the running CC model from session transcripts.
Exact numbers from pixels, for forensic recreation and for scoring a build: measure_frame(path, t, [crop, colors, ink_threshold]) returns the background hex, a k-means palette with shares, the ink mask's connected components as boxes in % of frame, row/column bands and their pitch (all at native resolution, so hexes are exact); measure_motion(path, from, to, [fps, crop, track]) returns hold runs and a step-Hz estimate, the list of state changes, luminance ramps, a radial profile and, with track: <hex>, colour-tracked keyframes with an ease hint. The same code runs from the shell as bun bin/lumiere-measure.mjs frame|motion for sessions where the MCP server is not connected.
Set persistent server defaults:
default_mode(low/mid/high/max): tier used when watch/measure omitmode.default_narrative_mode(true/false/"auto"): when true, narrative_mode defaults on for every watch/measure call that omits the param. When false, defaults off. When"auto"or unset, the heuristic decides per call (auto-suggest from motion / cuts / palette signals). Per-call param always wins.default_adaptive_sampling(true/false):trueactivates whenever motion_windows are cached and duration > 4s.falseforces uniform sampling. Default is off; opt-in only.backend(local/gemini-api/none): audio transcription backend.whisper_model(tiny/base/small/medium/large-v3-turbo/large-v3/auto): whisper size when backend=local.clear_sessions: true: wipe cached sessions + downloads.
Precedence for narrative_mode and adaptive_sampling: explicit per-call param > auto-suggest heuristic > server default > off.
| mode | px | best for |
|---|---|---|
low |
384 | overview, "what's this video about" |
mid |
512 | balanced reading, per-scene notes |
high (default) |
1024 | UI demos, animation, anatomy / feature reads |
max |
1536 | forensic detail, exact hex codes, sub-pixel events |
Higher tiers = fewer frames per call. For long action sequences at high or max, chunk by motion window.
watch with adaptive_sampling=true reads cached analyze.motion_windows and splits the frame budget 70/30 between motion and static spans. Motion windows get more frames per second (weighted by duration × intensity), static spans get spread thin. Same total frame count, but temporal resolution biased toward where action is happening.
When the moving subject is small relative to the frame (mascot inside a brand card), the subject's pixels get averaged into uniform color blobs at typical resolutions. ROI auto-crop reads analyze.subject_bbox (computed via connected-component segmentation of the binary motion mask) and crops watch frames to it before scaling. The subject fills the target resolution. 10x+ effective pixel density on the subject.
roi=auto uses ONE global bbox for the whole call, which has to be a union over every motion window's subject position. If the subject travels (mascot dashes top-left to bottom-right), that union balloons and the crop barely tightens. roi=per-window assigns each motion window its OWN bbox: analyze runs the cc-segmentation detector separately on each window's time range and stores them in analysis.window_bboxes. watch (with adaptive_sampling=true) then maps each adaptive segment to its window and crops with that window's bbox. Static segments fall back to the global subject_bbox. Result: max pixel density per window even when the subject moves across the frame.
Requires adaptive_sampling=true and a prior analyze(motion=true) call.
Frame-sampled perception has a structural failure: each frame gets interpreted in isolation, so continuous action reads as a sprite sheet of unrelated costumes. The fix is a temporal-narrative prompt:
- ANCHORS: identify what persists across frames
- CHANGES: identify what differs
- RESOLVE: each change as action / transition / state change
- NARRATE: as continuous prose
v0.12 uses one universal NARRATIVE_GUIDANCE prompt for all video types. The model determines what it's looking at from the frames alone, no domain-specific priors needed. Auto-suggested when prior analyze reports a content_class other than nature or generic.
analyze returns content_class as a structured 7-way enum for informational metadata. The classifier is signal-based, no extra ffmpeg passes, and uses a decision tree over motion summary (si / ti / subject ratio), scene cut density, palette outlier count, subject bbox method + confidence, and the transcription low-confidence flag. As of v0.12, content_class is no longer used to route narrative prompts but still drives auto-suggestion of narrative_mode.
Bbox detection runs a 3-tier cascade: connected-component segmentation (cleanest tight crop, confidence 0.9 when area_pct ∈ [10, 70]) → cropdetect (confidence 0.5 when area_pct < 85) → center-prior heuristic (60% center crop, confidence 0.2). The third tier is a v0.11 addition that gives roi=auto a useful crop on busy-background content (fitness videos, sports footage) where CC + cropdetect both fail.
When global motion windows cluster at the video boundaries (typical for fixed-camera footage with an off-center subject: the entry/exit walk registers as motion but the actual action does not), analyze runs a subject-region siti pass on the bbox crop. If that finds peaked windows, they replace the globals. If subject-region siti ALSO returns no peaks (typical for slow continuous motion like a deadlift or yoga flow), analyze synthesizes a single middle-60% best-guess action window so adaptive_sampling doesn't bias toward the boundary noise. motion_detection_warning surfaces the heuristic when it fires.
Score a render against a reference video, scene by scene, with only ffmpeg: block SSIM, ink SSIM (blocks that carry content in either frame, so a missing element shows), colour delta, worst frames, side-by-side sheets and red/green overlays, as report.md / report.json / report.html.
bun bin/lumiere-compare.mjs renders/zaro.mp4 ref/zaro-original.mp4 # scenes from the lock next to the render
bun bin/lumiere-compare.mjs renders/s04.mp4 ref/zaro-original.mp4 --scenes analysis/scenes.json # a windowed render: offset from its .window.json sidecar
bun bin/lumiere-compare.mjs out.mp4 ref.mp4 --scenes auto --gate 0.5 --json # scdet scenes, CI gate on ink SSIMA 70 s film at 10 fps scores in a few seconds (the sheets and overlays are the slow part, --no-sheets --overlays 0 for a bare score). Windowed renders from lumiere-render carry a <out>.window.json sidecar that sets the offset automatically. Samples sit on an absolute grid (t = k / fps), so a windowed render and a full render score the same instants.
Grammar mode (--mode grammar, the gate for a variant: the reference's structure with your own brand and copy) scores timing and motion instead of pixels: rhythm (diff-energy correlation), motion-event alignment within 100 ms, step match (held-frame ratio, so a 15 Hz film is recognised at a 30 fps grid), polarity-agnostic luminance ramps and ink-coverage correlation; --brand <brand.json> adds the on-brand pixel share and a reference-palette "reskin detector". Pixel SSIM is still reported, for information only.
bun bin/lumiere-compare.mjs renders/remit-variant.mp4 ../../zaro-launch/ref/zaro-original-4k.mp4 --mode grammar --brand brand.json --gate 0.6Measured on the Zaro subset: the 1:1 recreation scores grammar 0.79 with ref-palette 0.80 (a reskin signature by construction), the remit variant grammar 0.75 with on-brand 0.95 and ref-palette 0.00 (same beats, none of the reference's colours).
Standalone shell wrapper around the measure tool. Print exact tokens for a planned watch call without consuming them.
lumiere-cost /tmp/video.mp4 --mode high --end 00:00:24
lumiere-cost https://x.com/.../status/123 --mode max --start 00:00:05 --end 00:00:15
lumiere-cost video.mp4 --mode high --view-sample 50 --jsonRequires LUMIERE_ANTHROPIC_API_KEY in env or keychain.
Pick one:
# Option A: keychain (macOS)
security add-generic-password -a lumiere -s dev.lumiere-anthropic-api-key -w 'sk-ant-...'
# Option B: env
export LUMIERE_ANTHROPIC_API_KEY='sk-ant-...'The key is only used for /v1/messages/count_tokens (a free metering endpoint, no per-call charges).
The /lumiere skill instructs the model to:
inspect(path=URL)first- Read
cost_estimate.recommended_modeandcost_estimate.per_tier[current] - If full coverage would trigger autocompact, warn the user and offer (a) lower tier, (b)
/compactfirst, (c) a sub-segment via start_time/end_time analyzethenwatchwith appropriate chunks +narrative_modefor action segments- Optionally
measurefirst when the call is high-stakes (cap-tight, expensive, or budget-critical)
The SCAFFOLD phase: turn a storyboard into a rendered launch video.
/lumiere dashboard # open the studio (Library / Composer / Preview)
/lumiere create # scaffold from a storyboard or brief
The flow is author -> review in Preview -> export: write a freeform storyboard in the Composer (effect tokens + per-use overrides in parens), and a Claude session establishes a launch-video.lock.json (validated by creation/_tools/validate-lock.mjs) and restages each canonical effect into a composition per creation/RESTAGE.md. The composition is then reviewed in the dashboard Preview pane, which loads its registered window.__timelines.main and scrubs it frame-exactly (the same registry the renderer drives, so previewable == renderable). For multiple versions, point the Preview folder box at a directory holding a versions.json manifest and arrow prev/next across the set. Export to mp4 is the final step, once a version is chosen:
bun bin/lumiere-render.mjs <project-dir> # hyperframes CLI engine (default)
bun bin/lumiere-render.mjs <project-dir> --engine own # lumiere's own frame-exact pipeline
bun bin/lumiere-render.mjs <project-dir> --engine own-parallel # sharded fast 4K (--shards N, default 4)
bun bin/lumiere-render.mjs <project-dir> --engine own-parallel --resume # continue an interrupted render
bun bin/lumiere-render.mjs <project-dir> --engine own --scene s04 # v0.21: one lock scene only (+ <out>.window.json sidecar)
bun bin/lumiere-render.mjs <project-dir> --engine own --from 12 --to 14 # any time window (own / own-parallel)
bun bin/lumiere-render.mjs <project-dir> --engine own --still 9.5 --out renders/s.png # one PNG at tDefaults (fps, output, resolution) come from the lock. The own engine drives the registered paused timeline frame-by-frame via system Chrome + ffmpeg and matches the hyperframes frame count exactly; it needs Google Chrome installed and does not mix audio yet (use the default hyperframes engine for locked audio tracks). own-parallel shards that pipeline across N worker processes for fast 4K and is the engine for compositions that embed a per-frame-seeked <video>: it loads over file:// (so a file:// video src resolves, which the hyperframes http server blocks) and waits for each frame's seeked before capturing, so the video lands frame-exactly instead of tearing (re-encode embedded footage all-intra first, see RESTAGE "Embedded <video>"). Since v0.21.1 it is fast and crash-tolerant: captures use Chrome's fast PNG encoder (same pixels; --no-fast-png to opt out), --encoder auto switches to Apple VideoToolbox for 2160p+ output on Apple Silicon (--bitrate, default 45 Mbps; x264 crf 16 otherwise), and the frames are cut into --chunks (about 40 s each) drained by the --shards pool, so a crashed chunk retries once and an interrupted render continues with --resume. On the kint 4K film (296 s, 8896 frames) that took the render from 45+ minutes to 8. Every worker is a headless Google Chrome: quitting Chrome kills the chunks in flight. Both own engines are video-only; mux locked audio after with ffmpeg -i master.mp4 -i <audio> -map 0:v:0 -map 1:a:0 -c:v copy -c:a aac final.mp4.
Specs live in creation/ (LOCK.md, RESTAGE.md, the JSON schema) and effects/FORMAT.md (the per-effect file format + per-use ?vars= variables contract).
Recreate (/lumiere recreate <ref.mp4>): bin/lumiere-framebank.mjs writes a 10 fps frame bank with per-scene contact sheets; skills/lumiere/recreate.workflow.mjs runs Grammar -> Spec -> Verify (one analyst and one adversarial verifier per scene, one film-level grammar agent, every number from the measure tools); creation/_tools/collect-specs.mjs harvests finished scenes from the workflow journal into scene-spec.json. A lock that sets meta.reference and meta.sceneSpec and carries scenes[].measured restages in forensic mode (RESTAGE.md) and is gated by lumiere-compare.
Scene engines: each lock scene picks engine: dom | canvas | blender | ae. Canvas scenes paint on the timeline clock; Blender / After Effects scenes embed a PNG sequence (scenes[].source, validated by bin/lumiere-sequence.mjs) that the composition swaps per seek behind a decode gate the renderers await. The Preview scrubs every engine the same way, and shows the engine per scene.
Sub-compositions: a root index.html mounts one file per scene through data-composition-src with creation/_tools/subcomp-runtime.js (vendored into the project); children register paused timelines in local time, the root binds them to one master. A long film that tripped composition_file_too_large as one file lints clean split.
Brand layer: creation/BRAND.md + brand.schema.json; a lock binds a brand through meta.brand, brands can extend a parent (creation/brands/s0nderlabs.brand.json), and creation/_tools/resolve-brand.mjs merges the chain into the effective lock the validator checks.
Variant (/lumiere variant <ref.mp4> --brand <brand.json>): the reference's beat skeleton and motion grammar, your brand and content. scenes[].measured carries grammar only (step rate, holds, cadence, transitions, ease), never the reference's boxes or hexes; the gate is lumiere-compare --mode grammar --brand. The remit launch opening (~/Documents/s0nderlabs/remit-launch/variant-zaro/) is the worked example.
ElevenLabs SFX/BGM generation, own-engine audio mixing, and a formalized design-lock interview are not shipped yet. The INTERVIEW phase runs as a guided conversation today; the SCAFFOLD + RENDER phases are v0.16's ship.
MIT