Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Cutawan is a single **Electron + Vite + React + TypeScript** desktop app (npm, N
`xvfb-run -a --server-args="-screen 0 1600x1000x24" npx electron . --no-sandbox --disable-gpu`
(Requires a prior `npm run build` so `out/` exists.) `npm run dev` also works but expects a display.
- Harmless `Failed to connect to the bus` (DBus) and GPU warnings are expected under Xvfb and can be ignored.
- `scripts/smoke-test.sh [out-dir]` is the fastest end-to-end GUI check: it builds, seeds a demo project via `scripts/seed-demo.ts`, launches under Xvfb with `CUTAWAN_SMOKE` set, and writes `home.png`, `clips.png`, `editor.png`, `setup-clips.png`, `setup-caption-video.png`, `editor-caption-video.png`, `settings.png` and `settings-export.png` screenshots. The walk assumes a freshly seeded demo project (the script seeds one every run), and it runs "caption whole video" for real using the seeded transcript, so it makes no API calls. Set `CUTAWAN_SMOKE` to an output dir to trigger this auto-screenshot-and-exit mode.
- `scripts/smoke-test.sh [out-dir]` is the fastest end-to-end GUI check: it builds, seeds a demo project via `scripts/seed-demo.ts`, launches under Xvfb with `CUTAWAN_SMOKE` set, and writes the first-run wizard shots (`setup-wizard.png` plus `setup-openrouter*.png`, which read the sample catalogue in `tests/fixtures/openrouter-catalog.json` because the walk is offline), `home.png`, `clips.png`, `editor.png`, `setup-clips.png`, `setup-caption-video.png`, `editor-caption-video.png`, `settings.png` and `settings-export.png` screenshots. The walk assumes a freshly seeded demo project (the script seeds one every run), and it runs "caption whole video" for real using the seeded transcript, so it makes no API calls. Set `CUTAWAN_SMOKE` to an output dir to trigger this auto-screenshot-and-exit mode.

### Tests / lint
- `npm test` (vitest), `npm run typecheck` and `npm run lint` (eslint) are the static gates, and CI enforces all three on every push. Run them before proposing a change.
Expand Down
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,17 @@ the [releases page](https://github.com/JeremySNR/cutawan/releases).
This project uses [semantic versioning](https://semver.org/), loosely: while
still pre-1.0, minor bumps carry new features and patch bumps carry fixes.

## [Unreleased]

### Added

- **OpenRouter** is now a connection in first-run setup and Settings. Add an OpenRouter key (encrypted with your system keychain), then pick the clip-finding model from OpenRouter's full model list. The list is searchable, suggests models at the top, and shows prices, context size and image support. Transcription can use local Whisper on your computer or one of OpenRouter's Whisper models (Whisper, Whisper Large v3, Whisper Large v3 Turbo). These are the only OpenRouter transcription models that return the per-word timestamps captions need, so other transcription models aren't offered. **Check key** validates the key without making a model request.
- OpenRouter transcription sends five-minute audio chunks so each request finishes inside OpenRouter's 60-second limit. A reply without word timestamps fails with a message saying what to change, rather than producing captions that can't be timed.

### Improved

- The setup wizard no longer shows the unavailable "Claude subscription" card. Claude models can be used through OpenRouter instead.

## [0.13.0] - 2026-09-25

### Added
Expand Down
9 changes: 6 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@
## Download and get started

1. **[Download the latest release](https://github.com/JeremySNR/cutawan/releases/latest).** Choose the Windows `.exe` installer, macOS `.dmg`, or Linux `.AppImage` under **Assets**. You do not need Node.js or a source checkout to use the app.
2. **Set up your connection.** In v0.10.0 and newer, the first-run wizard offers **ChatGPT sign-in** via Codex for AI clip finding with local transcription, an **OpenAI-compatible API** for separately billed analysis, or **Local captions only** to caption a whole video without an AI connection. [Setup requirements and choices](docs/getting-started.md) are explained step by step. You can also explore the editor before setting up a connection.
2. **Set up your connection.** In v0.10.0 and newer, the first-run wizard offers **ChatGPT sign-in** via Codex for AI clip finding with local transcription, an **OpenAI-compatible API** for separately billed analysis, **OpenRouter** for any model OpenRouter lists with one key, or **Local captions only** to caption a whole video without an AI connection. [Setup requirements and choices](docs/getting-started.md) are explained step by step. You can also explore the editor before setting up a connection.
3. **Import a video or paste a supported URL.** Choose **Find viral clips** to review suggested moments, or **Caption whole video** to make one captioned edit. Adjust the trim, captions and framing, then export an MP4.

The ChatGPT/Codex option is a beta integration with your plan's Codex allowance, not an included OpenAI API. It needs the Codex CLI signed in with ChatGPT and Python 3.10+; the wizard can install local faster-whisper and a speech model after you choose it. AI clip finding still needs either Codex or an API connection. Builds before v0.10.0 do not show the wizard; configure the connection in **Settings → General → AI connection** instead.
Expand Down Expand Up @@ -217,6 +217,7 @@ for analysis. Rendering, face tracking, zoom and export are local.
Not as Cutawan's AI connection. [Anthropic's guidance](https://support.claude.com/en/articles/13189465-log-in-to-your-claude-account)
directs developers of third-party apps, including open-source apps, to use API-key
authentication. Cutawan will not route automated requests through a personal Claude login.
You can use Claude models with an API key through the **OpenRouter** connection instead.

**How is this different from Opus Clip's free tier?**
Free SaaS tiers cap your processing minutes and usually watermark the output.
Expand Down Expand Up @@ -260,8 +261,10 @@ Signing and notarisation are wanted; see
are configurable in Settings, including a cheaper legacy option.

**Can I run it against a local or non-OpenAI model?**
Yes, if it speaks the OpenAI REST shape. Set the API base URL in Settings
(or `OPENAI_BASE_URL`) to Azure OpenAI, OpenRouter, Groq, LM Studio, Ollama,
Yes. The **OpenRouter** connection takes one OpenRouter key and lets you pick
any chat model it lists, plus a hosted Whisper model or local transcription.
Your key is encrypted with the OS keychain. For anything else that speaks the OpenAI REST shape, set the API base URL in Settings
(or `OPENAI_BASE_URL`) to Azure OpenAI, Groq, LM Studio, Ollama,
or anything else with `/v1/chat/completions`. Transcription can point at a
separate local Whisper server (faster-whisper, whisper.cpp’s compatible
endpoint) — it must return **word-level timestamps**, because captions and
Expand Down
3 changes: 2 additions & 1 deletion docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,12 +18,13 @@ Signing and notarisation are planned. Until then, only download installers from

## Choose a connection

The first-run wizard offers three paths. You can switch later in **Settings → General → AI connection**. **Explore without setup** lets you look around, but processing a new video still needs one of the routes below.
The first-run wizard offers four paths. You can switch later in **Settings → General → AI connection**. **Explore without setup** lets you look around, but processing a new video still needs one of the routes below.

| Route | What it does | What you need |
| --- | --- | --- |
| **ChatGPT sign-in** (beta) | Uses Codex for clip finding and local Whisper for transcription. No OpenAI API key or automatic paid-API fallback. | An installed Codex CLI signed in with ChatGPT, Python 3.10+, and a local Whisper model. Your plan's Codex limits apply. [Detailed setup](chatgpt-subscription.md). |
| **OpenAI-compatible API** | Uses your configured API endpoint for analysis. Speech can go to the transcription API or run locally with Whisper. | An API key and separately billed provider usage. A local-speech choice also needs Python 3.10+ and a model. |
| **OpenRouter** | One OpenRouter key for clip finding with any chat model OpenRouter lists (OpenAI, Anthropic, Google and others), picked from a searchable list with suggestions at the top. Speech goes to an OpenRouter-hosted Whisper model or runs locally. | An [OpenRouter key](https://openrouter.ai/keys) and OpenRouter credits. A local-speech choice also needs Python 3.10+ and a model. |
| **Local captions only** | Transcribes and captions a whole video without sending it to an AI service. It does not find or score short clips. | Python 3.10+ and a local Whisper model. |

For either local-speech route, install Python separately first. The wizard can then create a private environment and download faster-whisper plus a **Small** or **Large v3** model into Cutawan's app-data folder after you request it. Small is the lighter download; Large v3 needs several gigabytes. The setup check does not make an AI model request. CPU transcription works but can be slow.
Expand Down
2 changes: 2 additions & 0 deletions scripts/smoke-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,9 @@ OUT="${1:-.tmp/smoke}"
mkdir -p "$OUT"
npm run build >/dev/null
SMOKE_OUT="$(realpath "$OUT")"
# The OpenRouter pickers read a saved sample catalogue so the walk needs no network.
CUTAWAN_USER_DATA="$SMOKE_OUT/wizard-profile" CUTAWAN_SMOKE="$SMOKE_OUT" CUTAWAN_SMOKE_WIZARD=1 \
CUTAWAN_OPENROUTER_CATALOG="$(realpath tests/fixtures/openrouter-catalog.json)" \
xvfb-run -a --server-args="-screen 0 1600x1000x24" \
npx electron . --no-sandbox --disable-gpu
npx tsx --tsconfig tsconfig.node.json scripts/seed-demo.ts
Expand Down
24 changes: 24 additions & 0 deletions src/main/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,30 @@ async function runSmokeCapture(win: BrowserWindow, dir: string): Promise<void> {
await sleep(2500)
if (process.env.CUTAWAN_SMOKE_WIZARD) {
await shot('setup-wizard')
// OpenRouter route: key entry plus the searchable model pickers.
const reveal = (selector: string): Promise<void> => win.webContents.executeJavaScript(
`document.querySelector(${JSON.stringify(selector)}).scrollIntoView({ block: 'start' })`
)
await click('[data-testid="setup-route-openrouter"]')
await reveal('[data-testid="openrouter-setup"]')
await shot('setup-openrouter')
await click('[data-testid="openrouter-model-toggle"]')
await reveal('[data-testid="openrouter-model"]')
await shot('setup-openrouter-models')
await win.webContents.executeJavaScript(`(() => {
const input = document.querySelector('[data-testid="openrouter-model-search"]')
Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, 'value').set.call(input, 'claude')
input.dispatchEvent(new Event('input', { bubbles: true }))
})()`)
await sleep(400)
await shot('setup-openrouter-search')
await click('[data-testid="openrouter-model-toggle"]')
await click('[data-testid="openrouter-transcription-toggle"]')
await reveal('[data-testid="openrouter-transcription"]')
await shot('setup-openrouter-transcription')
await click('[data-testid="openrouter-transcription-option-local-whisper"]')
await reveal('[data-testid="openrouter-transcription"]')
await shot('setup-openrouter-local')
app.quit()
return
}
Expand Down
6 changes: 5 additions & 1 deletion src/main/ipc.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import { checkLocalWhisperSetup, checkSubscriptionSetup } from './subscription'
import { checkOpenRouterKey, listOpenRouterModels } from './openrouter'
import { cancelLocalWhisperInstall, installLocalWhisper, type LocalWhisperModel } from './localWhisper'
import { app, BrowserWindow, dialog, ipcMain, shell } from 'electron'
import { existsSync } from 'node:fs'
Expand Down Expand Up @@ -38,6 +39,7 @@ import {
getExportPreferences,
getModelPreferences,
getSettings,
missingCredentialName,
updateSettings
} from './settings'

Expand Down Expand Up @@ -332,7 +334,7 @@ export function registerIpcHandlers(): void {
const clip = project.clips.find((c) => c.id === clipId)
if (!clip) throw new Error('Clip not found')
const apiKey = getAnalysisCredential()
if (!apiKey) throw new Error('Add your OpenAI API key in Settings first.')
if (!apiKey) throw new Error(`Add your ${missingCredentialName()} in Settings first.`)
const caption = await generateSocialCaption(
apiKey,
getModelPreferences().analysisModel,
Expand Down Expand Up @@ -363,6 +365,8 @@ export function registerIpcHandlers(): void {

handle('settings:checkSubscription', () => checkSubscriptionSetup())
handle('settings:checkLocalWhisper', () => checkLocalWhisperSetup())
handle('openrouter:models', (_e, refresh?: boolean) => listOpenRouterModels(refresh === true))
Comment thread
JeremySNR marked this conversation as resolved.
handle('openrouter:checkKey', (_e, key?: string) => checkOpenRouterKey(key))
handle('settings:installLocalWhisper', async (event, model: LocalWhisperModel, pythonPath: string) => {
const result = await installLocalWhisper(model, pythonPath, progress => {
if (!event.sender.isDestroyed()) event.sender.send('whisper:installProgress', progress)
Expand Down
94 changes: 94 additions & 0 deletions src/main/openrouter.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
import { readFile } from 'node:fs/promises'
import {
OPENROUTER_API_BASE,
parseOpenRouterModels,
suggestedAsModels,
SUGGESTED_OPENROUTER_MODELS,
OPENROUTER_TRANSCRIPTION_MODELS,
isSupportedTranscriptionModel,
type OpenRouterCatalog
} from '@shared/openrouter'
import { getOpenRouterKey } from './settings'

const CATALOG_TTL_MS = 10 * 60_000
const REQUEST_TIMEOUT_MS = 12_000

let cached: { at: number; catalog: OpenRouterCatalog } | null = null

async function getJson(path: string, key?: string): Promise<{ status: number; body: unknown }> {
const res = await fetch(`${OPENROUTER_API_BASE}${path}`, {
headers: key ? { Authorization: `Bearer ${key}` } : {},
signal: AbortSignal.timeout(REQUEST_TIMEOUT_MS)
})
let body: unknown = null
try { body = await res.json() } catch { /* non-JSON body */ }
return { status: res.status, body }
}

/**
* The public model catalogue (no key needed): chat models, and the
* transcription models OpenRouter serves at /audio/transcriptions that
* Cutawan can use (word timestamps required). Offline,
* the suggested models stand in so setup still works.
* `CUTAWAN_OPENROUTER_CATALOG` points at a saved `/models` response, for
* headless checks without network access.
*/
export async function listOpenRouterModels(refresh = false): Promise<OpenRouterCatalog> {
if (!refresh && cached && Date.now() - cached.at < CATALOG_TTL_MS) return cached.catalog
const fixture = process.env.CUTAWAN_OPENROUTER_CATALOG
try {
let llm, transcription
if (fixture) {
const saved = JSON.parse(await readFile(fixture, 'utf8')) as { models: unknown; transcription: unknown }
llm = parseOpenRouterModels(saved.models, 'llm')
transcription = parseOpenRouterModels(saved.transcription, 'transcription')
} else {
const [chat, speech] = await Promise.all([
getJson('/models'),
getJson('/models?output_modalities=transcription').catch(() => ({ status: 0, body: null }))
])
if (chat.status !== 200) throw new Error(`OpenRouter returned HTTP ${chat.status}`)
llm = parseOpenRouterModels(chat.body, 'llm')
transcription = parseOpenRouterModels(speech.body, 'transcription')
}
if (!llm.length) throw new Error('OpenRouter returned no models')
// Only models that return word timestamps; see OPENROUTER_TRANSCRIPTION_MODELS.
// A short or failed transcription listing keeps the whole allowlist.
const supported = transcription.filter(m => isSupportedTranscriptionModel(m.id))
const catalog: OpenRouterCatalog = {
llm,
transcription: supported.length ? supported : suggestedAsModels(OPENROUTER_TRANSCRIPTION_MODELS),
live: true
}
cached = { at: Date.now(), catalog }
return catalog
} catch (error) {
return {
llm: suggestedAsModels(SUGGESTED_OPENROUTER_MODELS),
transcription: suggestedAsModels(OPENROUTER_TRANSCRIPTION_MODELS),
live: false,
error: error instanceof Error ? error.message : String(error)
}
}
}

/**
* Validate a key (the typed one, else the stored one) with OpenRouter's
* key-info endpoint. Makes no model request and costs nothing.
*/
export async function checkOpenRouterKey(typed?: string): Promise<{ message: string }> {
const key = typed?.trim() || getOpenRouterKey()
if (!key) throw new Error('Enter an OpenRouter API key first.')
let res: { status: number; body: unknown }
try {
res = await getJson('/key', key)
} catch (error) {
throw new Error(`Could not reach OpenRouter: ${error instanceof Error ? error.message : String(error)}`, { cause: error })
}
if (res.status === 401 || res.status === 403) throw new Error('OpenRouter rejected this key. Copy it again from openrouter.ai/keys.')
if (res.status !== 200) throw new Error(`OpenRouter key check failed (HTTP ${res.status}).`)
const data = (res.body as { data?: { limit_remaining?: unknown; is_free_tier?: unknown } } | null)?.data
const remaining = typeof data?.limit_remaining === 'number' ? ` $${data.limit_remaining.toFixed(2)} of this key’s limit remains.` : ''
const free = data?.is_free_tier === true ? ' This account is on the free tier; add credits to use paid models.' : ''
return { message: `OpenRouter accepted the key.${remaining}${free}` }
}
16 changes: 9 additions & 7 deletions src/main/pipeline/ffmpeg.ts
Original file line number Diff line number Diff line change
Expand Up @@ -220,7 +220,8 @@ export interface AudioChunk {
* so 20-minute chunks (~7.2 MB) leave plenty of headroom. Consecutive chunks
* overlap by a few seconds so no word is cut in half at a boundary — Whisper
* sees full context on both sides and the stitcher picks each word from the
* chunk that owns its timestamp.
* chunk that owns its timestamp. Routes with a short request timeout pass a
* smaller `chunkSec`.
*/
export const AUDIO_CHUNK_SEC = 20 * 60
export const AUDIO_CHUNK_OVERLAP_SEC = 8
Expand All @@ -230,27 +231,28 @@ export async function extractAudioChunks(
workDir: string,
durationSec: number,
onProgress?: (fraction: number) => void,
signal?: AbortSignal
signal?: AbortSignal,
chunkSec = AUDIO_CHUNK_SEC
): Promise<AudioChunk[]> {
await mkdir(workDir, { recursive: true })
const stride = AUDIO_CHUNK_SEC - AUDIO_CHUNK_OVERLAP_SEC
const stride = chunkSec - AUDIO_CHUNK_OVERLAP_SEC
const count =
durationSec <= AUDIO_CHUNK_SEC ? 1 : 1 + Math.ceil((durationSec - AUDIO_CHUNK_SEC) / stride)
durationSec <= chunkSec ? 1 : 1 + Math.ceil((durationSec - chunkSec) / stride)
const chunks: AudioChunk[] = []
for (let i = 0; i < count; i++) {
const offset = i * stride
const out = join(workDir, `audio-${i}.mp3`)
const args = [
'-ss', String(offset),
'-t', String(AUDIO_CHUNK_SEC),
'-t', String(chunkSec),
'-i', videoPath,
'-vn',
'-ac', '1',
'-ar', '16000',
'-b:a', '48k',
out
]
const chunkDur = Math.min(AUDIO_CHUNK_SEC, durationSec - offset)
const chunkDur = Math.min(chunkSec, durationSec - offset)
await runFfmpeg(args, {
onProgress: (t) => onProgress?.(Math.min(1, (offset + Math.min(t, chunkDur)) / durationSec)),
signal
Expand All @@ -260,7 +262,7 @@ export async function extractAudioChunks(
offsetSec: offset,
keepFromSec: i === 0 ? 0 : offset + AUDIO_CHUNK_OVERLAP_SEC / 2,
keepToSec:
i === count - 1 ? Number.POSITIVE_INFINITY : offset + AUDIO_CHUNK_SEC - AUDIO_CHUNK_OVERLAP_SEC / 2
i === count - 1 ? Number.POSITIVE_INFINITY : offset + chunkSec - AUDIO_CHUNK_OVERLAP_SEC / 2
})
}
onProgress?.(1)
Expand Down
4 changes: 2 additions & 2 deletions src/main/pipeline/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ import {
YtDlpError,
type CookieAuthOptions
} from './ytdlp'
import { getAnalysisCredential, getImportPreferences, getModelPreferences } from '../settings'
import { getAnalysisCredential, getImportPreferences, getModelPreferences, missingCredentialName } from '../settings'
import { projectDir, saveProject, updateProject } from '../projects'

export async function createProject(videoPath: string): Promise<Project> {
Expand Down Expand Up @@ -142,7 +142,7 @@ export async function analyzeProject(
): Promise<Project> {
const apiKey = getAnalysisCredential()
if (!apiKey) {
throw new Error('No API key configured. Add one in Settings before generating clips.')
throw new Error(`No ${missingCredentialName()} configured. Add one in Settings before generating clips.`)
}
const settings = getModelPreferences()
const generationId = randomUUID()
Expand Down
Loading
Loading