Which models VibeFrame calls, how to pick one, and the provider behaviour worth knowing.
The model list itself is generated from the model catalog (packages/ai-providers/src/catalog/catalog.ts), which the CLI also reads for aliases, defaults, and cost estimates.
Generated from the model catalog; edit packages/ai-providers/src/catalog/catalog.ts, then run pnpm gen:models.
Prices are upper-bound list prices for estimates, not metered billing.
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| claude | Claude Sonnet 5.5 (default) | claude-sonnet-5-5 |
sonnet |
GA | $2 / 1M input tokens | Agent, storyboard, translate-srt, motion default |
| claude | Claude Opus 5.5 | claude-opus-5-5 |
opus |
GA | $4 / 1M input tokens | |
| claude | Claude Opus 4.6 | claude-opus-4-6 |
opus-4-6 |
Legacy | $5 / 1M input tokens | |
| openai | GPT-5.4 mini (default) | gpt-5.4-mini |
GA | $0.75 / 1M input tokens | Agent and storyboard default (tools on Chat Completions) | |
| gemini | Gemini 3.8 Flash (default) | gemini-3.8-flash |
flash, latest, flash-3.8, gemini |
GA | $0.75 / 1M input tokens (until 2026-12-31, then $1.50) | |
| gemini | Gemini 3.5 Flash | gemini-3.5-flash |
flash-3.5 |
GA | ||
| gemini | Gemini 3 Flash (preview) | gemini-3-flash-preview |
flash-3 |
Preview | ||
| gemini | Gemini 3.1 Pro (preview) | gemini-3.1-pro-preview |
pro, pro-3.1, 3.1-pro, gemini-3.1-pro |
Preview | ||
| gemini | Gemini 2.5 Flash | gemini-2.5-flash |
flash-2.5 |
Legacy | New projects cannot use 2.5 models since 2026-09-18 | |
| gemini | Gemini 2.5 Pro | gemini-2.5-pro |
pro-2.5, 2.5-pro |
Legacy | ||
| grok | Grok 4.3 (default) | grok-4.3 |
GA | $1.25 / 1M input tokens | ||
| grok | Grok 4.7 | grok-4.7 |
GA | $2 / 1M input tokens |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| openai | GPT Image 2.5 Sunburst (default) | gpt-image-2.5-sunburst |
2.5, sunburst |
GA | $0.211 / image (high, 1024x1024) | |
| openai | GPT Image 2.5 Flare | gpt-image-2.5-flare |
flare, 2.5-flare |
GA | $0.211 / image (high, 1024x1024) | |
| openai | GPT Image 2 | gpt-image-2 |
2 |
GA | $0.211 / image (high, 1024x1024) | |
| openai | GPT Image 1.5 | gpt-image-1.5 |
1.5 |
Deprecated, shuts down 2026-12-01 | $0.133 / image (high, 1024x1024) | |
| gemini | Nano Banana 2 (default) | gemini-3.1-flash-image |
flash, 3.1-flash, latest |
GA | $0.067 / image (1K) | |
| gemini | Nano Banana 2 Lite | gemini-3.1-flash-lite-image |
lite |
GA | $0.034 / image (1K) | |
| gemini | Nano Banana Pro | gemini-3-pro-image |
pro |
GA | $0.134 / image (1K-2K) | |
| gemini | Nano Banana (2.5 Flash Image) | gemini-2.5-flash-image |
Deprecated, shuts down 2026-10-02 | |||
| gemini | Nano Banana 2 (preview) | gemini-3.1-flash-image-preview |
Deprecated, shuts down 2026-06-25 | |||
| gemini | Nano Banana Pro (preview) | gemini-3-pro-image-preview |
Deprecated, shuts down 2026-06-25 | |||
| grok | Grok Imagine Image (default) | grok-imagine-image |
GA | $0.02 / image | ||
| grok | Grok Imagine Image 2.0 | grok-imagine-image-2.0 |
pro, 2.0, quality |
GA | $0.08 / image (medium, 2K) |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| seedance | Seedance 2.0 (default) | seedance-2.0 |
quality, 2.0 |
GA | $0.3024 / second (720p 16:9) | 4-15 s, native audio |
| seedance | Seedance 2.0 Fast | seedance-2.0-fast |
fast |
GA | $0.2419 / second (720p 16:9) | 4-15 s, up to 720p, native audio |
| seedance | Seedance 2.5 | seedance-2.5 |
2.5 |
GA | $0.473 / second (720p 16:9) | Opt-in; 4-30 s, native audio |
| omni | Gemini Omni 1.1 Flash (default) | gemini-omni-1.1-flash |
GA | $0.1 / second (720p) | Google video default; the model picks 3-10 s; native audio | |
| veo | Veo 3.1 | veo-3.1-generate-preview |
3.1 |
Deprecated, shuts down 2026-10-22; use -p omni (Gemini Omni 1.1 Flash) |
$0.4 / second (720p/1080p) | 4-8 s, native audio; explicit -p veo only |
| veo | Veo 3.1 Fast (default) | veo-3.1-fast-generate-preview |
3.1-fast |
Deprecated, shuts down 2026-10-22; use -p omni (Gemini Omni 1.1 Flash) |
$0.1 / second (720p) | 4-8 s, native audio; explicit -p veo only |
| grok | Grok Imagine Video 1.5 (default) | grok-imagine-video-1.5 |
1.5 |
GA | $0.14 / second (720p) | 1-15 s, native audio; 720p unless --resolution says otherwise |
| grok | Grok Imagine Video 1.5 Lite | grok-imagine-video-1.5-lite |
lite, 1.5-lite |
GA | $0.03 / second (720p) | 1-15 s, native audio |
| grok | Grok Imagine Video | grok-imagine-video |
classic |
Legacy | $0.05 / second (480p) | Previous generation; the only Grok model that edits and extends videos |
| kling | Kling v3 (default) | kling-v3 |
v3 |
GA | $0.084 / second (std, silent) | 3-15 s, multi-shot; std 720p, pro 1080p |
| kling | Kling v2.6 | kling-v2-6 |
v2.6 |
GA | $0.042 / second (std 720p) | 5 or 10 s |
| kling | Kling v2.5 Turbo | kling-v2-5-turbo |
v2.5-turbo, v2.5 |
GA | $0.042 / second (std) | 5 or 10 s, no audio |
| runway | Runway Gen-4.5 (default) | gen4.5 |
GA | $0.12 / second | 2-10 s, no audio | |
| runway | Runway Gen-4 Turbo | gen4_turbo |
GA | $0.05 / second | Image-to-video only, no audio |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| elevenlabs | ElevenLabs v3 (default) | eleven_v3 |
v3 |
GA | $0.08 / 1K characters | 5,000 characters per request |
| elevenlabs | ElevenLabs Multilingual v2 | eleven_multilingual_v2 |
multilingual, v2 |
GA | $0.08 / 1K characters | 10,000 characters; steadiest on long form |
| elevenlabs | ElevenLabs Flash v2.5 | eleven_flash_v2_5 |
flash |
GA | $0.04 / 1K characters | 40,000 characters; lowest latency |
| openai | GPT-4o mini TTS (default) | gpt-4o-mini-tts |
Deprecated, shuts down 2027-01-06 | |||
| openai | TTS-1 | tts-1 |
Deprecated, shuts down 2027-01-06 | $0.015 / 1K characters | ||
| openai | TTS-1 HD | tts-1-hd |
Deprecated, shuts down 2027-01-06 | $0.03 / 1K characters | ||
| kokoro | Kokoro 82M (default) | onnx-community/Kokoro-82M-v1.0-ONNX |
kokoro |
GA | $0 / 1K characters | Runs locally; ~90 MB download on first use |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| elevenlabs | ElevenLabs Music v2.5 (default) | music_v2_5 |
GA | $0.15 / minute |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| elevenlabs | ElevenLabs Sound Effects v2 (default) | eleven_text_to_sound_v2 |
GA | $0.12 / minute | 0.5-30 s |
| Provider | Model | ID | Aliases | Status | Price | Notes |
|---|---|---|---|---|---|---|
| openai | Whisper (default) | whisper-1 |
Deprecated, shuts down 2027-02-26 | Only OpenAI model with word timestamps |
| Command | Flag | Values |
|---|---|---|
vibe agent -p claude|openai|gemini|xai |
--model |
Any model ID; default is the provider's catalog default |
vibe generate motion |
-m |
Claude aliases (sonnet default, opus, opus-4-6) or Gemini aliases (gemini, gemini-3.1-pro, gemini-2.5-pro) |
vibe generate image -p openai|gemini|grok, vibe edit image |
-m |
The image aliases above |
vibe generate video -p seedance |
--seedance-model |
quality (default), fast, 2.5 |
vibe generate video -p grok |
--grok-model |
1.5 (default), lite, classic |
vibe generate video -p kling |
--kling-model |
v3 (default), v2.6, v2.5-turbo |
vibe generate video -p veo |
--veo-model |
3.1-fast (default), 3.1 |
vibe generate video -p runway |
--runway-model |
gen4.5 (default), gen4_turbo |
Used by the built-in agent mode (vibe agent), which is deprecated and will be removed before 1.0.
Let your coding agent or an MCP host drive vibe instead.
| Provider | Model | API Model ID | Env Key | CLI Option |
|---|---|---|---|---|
| OpenAI | GPT-5.4 mini | gpt-5.4-mini |
OPENAI_API_KEY |
-p openai |
| Claude | Sonnet 5.5 | claude-sonnet-5-5 |
ANTHROPIC_API_KEY |
-p claude |
| Gemini | 3.8 Flash | gemini-3.8-flash |
GOOGLE_API_KEY |
-p gemini |
| xAI | Grok 4.3 | grok-4.3 |
XAI_API_KEY |
-p xai |
| OpenRouter | Auto (300+ models) | openrouter/auto |
OPENROUTER_API_KEY |
-p openrouter |
| Evolink | GPT-5.2 (via unified API) | gpt-5.2 |
EVOLINK_API_KEY |
-p evolink |
| Ollama | Local models | user-configured | - | -p ollama |
Override the model per session with vibe agent -p <provider> --model <id>.
Claude and Gemini replay their own assistant content (thinking blocks, thought signatures) between tool turns, which current models require.
Claude Opus 4.7 and every 5.x model reject temperature, top_p, and top_k, and Gemini 3.x deprecates them; VibeFrame never sends them.
GPT-6 Astra and 6.1 Sol reject tool calls on Chat Completions (Responses API only), so they cannot drive Agent mode yet.
grok-4-1-fast-reasoning was retired on 2026-05-15; xAI silently serves grok-4.3 for it.
OpenRouter model options:
openrouter/auto is the default and routes to a model per request. OpenRouter slugs use dots in version numbers:
| Model ID | Provider | Notes |
|---|---|---|
openrouter/auto |
Auto | Default. Cost and tool support vary per pick |
anthropic/claude-sonnet-5.5 |
Anthropic | Claude Sonnet 5.5 |
openai/gpt-6-luna |
OpenAI | GPT-6 Luna |
google/gemini-3.8-flash |
Gemini 3.8 Flash | |
qwen/qwen3.8-flash |
Qwen | Inexpensive, tool-capable |
deepseek/deepseek-v4.1-flash |
DeepSeek | Inexpensive, tool-capable |
To use a specific model: vibe agent -p openrouter --model anthropic/claude-sonnet-5.5
See openrouter.ai/models for the full list.
Evolink model options:
| Model ID | Provider | Notes |
|---|---|---|
gpt-5.2 |
OpenAI | Default. GPT-5.2 via Evolink |
evolink/auto |
Auto | Auto-routing |
gemini-3.8-flash |
Tool calls work on Chat Completions | |
claude-sonnet-5-5 |
Anthropic | Claude Sonnet 5.5 via Evolink |
doubao-seed-2.0-pro |
ByteDance | Doubao Seed 2.0 Pro via Evolink |
See docs.evolink.ai for the full model catalog.
Remotion component code generation takes a Claude or Gemini alias with -m (see "Choosing a model").
Claude motion replies use structured outputs, so the generated component code always arrives as valid JSON.
Models, aliases, and prices are in the catalog above; -p openai|gemini|grok picks the provider.
Images are saved in the format the output file name asks for: when a provider returns JPEG for a .png path, VibeFrame converts it with FFmpeg.
gemini-2.5-flash-image and the -preview Nano Banana IDs are past their shutdown dates; pass them explicitly only if you must.
All Gemini image models support 14 aspect ratios: 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9
Grok Imagine supports 14 aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 2:1, 1:2, 19.5:9, 9:19.5, 20:9, 9:20, auto
| Provider | Model | Max Input Images | CLI Option | Features |
|---|---|---|---|---|
| Gemini | gemini-3.1-flash-image |
3 | -p gemini (default) |
Fast editing |
| Gemini | gemini-3-pro-image |
14 | -p gemini -m pro |
Multi-image composition, up to 4K output |
| OpenAI | gpt-image-2.5-sunburst |
16 | -p openai |
Instruction-based editing, multipart upload |
| xAI Grok | grok-imagine-image |
1 | -p grok |
Single image editing, $0.02/edit; -m pro uses grok-imagine-image-2.0 |
Models marked Audio: Yes generate synchronized sound (dialogue, SFX, ambient). Silent models need separate
vibe generate speech/vibe generate sound-effect.
Models, durations, and prices are in the catalog above.
Provider keys: Seedance FAL_API_KEY, Grok XAI_API_KEY, Kling KLING_API_KEY, Omni and Veo GOOGLE_API_KEY, Runway RUNWAY_API_SECRET.
-p falis a deprecated v0.x alias for-p seedanceand will be removed at the 1.0 cut. Use-p seedancein new scripts.
⚠️ Seedance rejects image-to-video inputs showing a recognizable face. This is ByteDance's platform-wide likeness policy, not a fal.ai quirk: the API returns a deterministic HTTP 422 ("The images or videos provided may contain likenesses of real people") for keyframes where a person's face is clearly visible - including AI-generated photoreal faces. Hands-only or back-of-head shots pass. Retrying does not help.vibe buildfalls back to Runway automatically for such beats whenRUNWAY_API_SECRETis configured; you can also pin a beat with aprovider: runwaycue, or runvibe generate video "<motion>" -p runway -i keyframe.pngdirectly. Observed 2026-07-26 onseedance-2.0.
Veo 3.1 previews shut down on 2026-10-22. Google names Gemini Omni 1.1 Flash as the replacement, so Omni is now the Google video default (
-p omni, or auto-selected whenGOOGLE_API_KEYis the only video key) andvibe build --video-provider omniworks.-p veostill works until the shutdown and prints a warning; from 2026-10-22 it fails before any request and points to-p omni.
| Option | Values | Description |
|---|---|---|
--negative-prompt |
text | What to avoid in the generated video |
--resolution |
720p, 1080p, 4k | Video resolution |
--last-frame |
image path | Frame interpolation (first→last frame) |
--ref-images |
image paths (max 3) | Character consistency (Veo 3.1 only) |
--person |
allow_all, allow_adult | Person generation setting |
veo-extend |
operation-name | Extend a previously generated Veo video |
All text-to-video providers also support image-to-video. Key differences per provider:
| Provider | Model | I2V Support | Image Input | Notes |
|---|---|---|---|---|
| Seedance via fal.ai | seedance-2.0 |
Yes | URL only | Auto-uploads via ImgBB (IMGBB_API_KEY) for local image paths. |
| xAI Grok | grok-imagine-video-1.5 |
Yes | URL or data URI | Same pricing as T2V, plus $0.01 per input image |
| Kling | all v2.5+ models | Yes | URL only | Auto-uploads via ImgBB (IMGBB_API_KEY) |
| Veo | all models | Yes | base64 (first frame) | Supports --last-frame for frame interpolation |
| Runway | gen4.5 |
Yes | URL or data URI | Text+image-to-video |
| Runway | gen4_turbo |
I2V only | URL or data URI | Cannot do text-only generation |
| Gemini Omni | gemini-omni-1.1-flash |
Yes | base64 (first frame; --last-frame adds an end frame) |
Sent as typed image inputs before the text prompt |
Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) is Google's GA video model on the Generative Language API and the replacement for the Veo 3.1 previews.
Reference: https://ai.google.dev/gemini-api/docs/omni
| Property | Value |
|---|---|
| Model ID | gemini-omni-1.1-flash |
| Endpoint | POST /v1beta/interactions (not Veo's :predictLongRunning) |
| Auth | GOOGLE_API_KEY (same key as Gemini) |
| Input | text, or typed image items (first and last frame) followed by the text |
| Output | response_format: {type: "video", aspect_ratio, resolution, delivery: "uri"}; VibeFrame waits for the Files API entry to become ACTIVE, then downloads it with the key in a header |
| Aspect ratios | 16:9, 9:16 |
| Resolution | 360p, 720p (default); 1080p and 4k are upscaled |
| Duration | 3-10 s, chosen by the model (no duration parameter) |
| Watermark | SynthID |
Not yet wired: multi-turn editing with previous_interaction_id, extension up to 40 s, and video references.
Recognizable real people in uploaded images are not supported, so it is no fallback for Seedance's likeness filter.
| Provider | Capability | Env Key | Notes |
|---|---|---|---|
| ElevenLabs | TTS, SFX, Music, Voice Clone | ELEVENLABS_API_KEY |
TTS: eleven_v3. SFX: 0.5-30s. Music: 3s-10min, model music_v2_5 (48 kHz / 192 kbps MP3) |
| OpenAI TTS | TTS | OPENAI_API_KEY |
gpt-4o-mini-tts (~$0.015/min of audio); voices incl. marin, alloy, nova |
| Kokoro | TTS (local, free) | - | Kokoro-82M (Apache 2.0); ~90MB model on first use; bundled in the Desktop extension. English voices only. Long narration is synthesised in sentence chunks |
| Whisper | Transcription | OPENAI_API_KEY |
OpenAI API |
| Replicate | Music generation | REPLICATE_API_TOKEN |
MusicGen, max 30s. Weights are CC-BY-NC: not for commercial use |
# LLM Providers
export OPENAI_API_KEY="sk-..." # GPT, Whisper, GPT Image 2.5
export ANTHROPIC_API_KEY="sk-ant-..." # Claude
export GOOGLE_API_KEY="AIza..." # Gemini (image, Veo video)
export XAI_API_KEY="xai-..." # xAI Grok
export OPENROUTER_API_KEY="sk-or-..." # OpenRouter (300+ models)
export EVOLINK_API_KEY="el-..." # Evolink (GPT-5, Claude, Gemini, DeepSeek & more)
# Media Providers
export ELEVENLABS_API_KEY="..." # TTS, SFX
export RUNWAY_API_SECRET="..." # Runway Gen-4 Turbo
export KLING_API_KEY="..." # Kling v2.x/v3
export REPLICATE_API_TOKEN="..." # Replicate (music)| Command | Required API Key | Model |
|---|---|---|
vibe agent (default) |
OPENAI_API_KEY |
GPT-5.4 mini (Agent LLM) |
vibe agent -p claude |
ANTHROPIC_API_KEY |
Claude Sonnet 5.5 (Agent LLM) |
vibe agent -p gemini |
GOOGLE_API_KEY |
Gemini 3.8 Flash (Agent LLM) |
vibe agent -p xai |
XAI_API_KEY |
Grok 4.3 (Agent LLM) |
vibe agent -p openrouter |
OPENROUTER_API_KEY |
OpenRouter Auto (Agent LLM) |
vibe agent -p evolink |
EVOLINK_API_KEY |
GPT-5.2 via Evolink (Agent LLM) |
vibe generate image -p openai |
OPENAI_API_KEY |
OpenAI image generation |
vibe generate image -p gemini |
GOOGLE_API_KEY |
Gemini image generation |
vibe edit image |
GOOGLE_API_KEY |
Gemini Nano Banana |
vibe generate speech |
ELEVENLABS_API_KEY |
ElevenLabs |
vibe generate music |
ELEVENLABS_API_KEY |
ElevenLabs Music (default) |
vibe generate music -p replicate |
REPLICATE_API_TOKEN |
Replicate MusicGen |
vibe generate video -p seedance |
FAL_API_KEY |
Seedance via fal.ai |
vibe generate video -p grok |
XAI_API_KEY |
Grok Imagine Video 1.5 |
vibe generate video -p kling |
KLING_API_KEY |
Kling v3 |
vibe generate image -p grok |
XAI_API_KEY |
Grok Imagine |
vibe generate video -p omni |
GOOGLE_API_KEY |
Gemini Omni 1.1 Flash |
vibe generate video -p veo |
GOOGLE_API_KEY |
Veo 3.1 (shuts down 2026-10-22) |
Provider defaults depend on which API keys are configured. For public docs and demos, prefer explicit provider flags so the required key is obvious:
vibe generate image "..." -p openai
vibe generate image "..." -p gemini
vibe generate video "..." -p seedance
vibe generate video "..." -p veoUse command help as the runtime source of truth for supported flags:
vibe generate image --help
vibe generate video --help
vibe edit image --help