Repository navigation
Replies: 3 comments 2 replies
10-20 minutes is really poor performance, This is definitely an area we can improve in. I'll look into more transcription engines and how we can improve the Transcription tool and add more features/options like selecting between per word transcription and other alternatives. |
Concat does have an MCP (Settings > Remote) that you can connect to. I had Claude generate some quick docs for those who are interested, though I haven't had the time to evaluate and test the docs myself. |
|
Another thing I remembered is that Parakeet on CPU itself is EXTREMELY fast compared to Whisper (on CPU). So even just adding parakeet support to the model list to the next release would speed things up without the effort needed to test GPU (mappings, binaries and whatever else). |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Currently transcription (for subtitle generation) seems to be whisper on CPU only, which takes a VERY long time for clips that are 10/20m long.
Is it possible to add support whisper on GPU, as well as other engines like WhisperX, faster-whisper (CTranslate2), parakeet (NVIDIA), Sherpa-ONNX, etc. on the same machine?
There is also no per-word transcription - I only get clusters of words transcribed per one caption block. An engine that does per-word transcription would make agentic editing even faster (more accurate at tasks like aligning clips on different tracks based on their transcript).
Also, is it possible to add support for remote whisper transcription (e.g. Groq API which has a VERY generous free tier, OpenAI's API themselves, or even selfhosted (this would mean following verbose_json spec, Speaches project is a good example, it follows OpenAI audio API & verbose_json standard)
All reactions