You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Thanks for making this, it's so nice to be able to sample all the TTSes in one place without having to get the models. It's appreciated. I have an ARM system and almost no TTSes work out of the box.
Some random thoughts:
A new TTS just came out, dots.tts 2B. If the demos aren't cherry-picked this is the best TTS I've ever heard, including best voice cloning and best expressivity. You really want to add this. https://rednote-hilab.github.io/dots.tts-demo/
Could you add MOSS TTS 1.5 which came out last month? It's not as impressive but it's pretty solid.
Why are there only partial entries on some prompts? For example prompt 5 has way fewer entries than prompt 3. Does it take a long time to do a full run with all supported models?
Is there some source that tracks WER for all models on the same benchmark? And if so, would you want to add it in a column on the Listen page, next to the Play Audio button for each model? This way, it helps quickly see the trade-offs of a great-sounding model. For example your 'TLDR best cloning choice,' OmniVoice, on Prompt 3 voice clone skips the entire beginning of the sentence (5 words). It feels wrong to make this your Great comparison #1 recommendation when it has such a glaring flaw, I don't care how good the voice is.
Thanks for making this, it's so nice to be able to sample all the TTSes in one place without having to get the models. It's appreciated. I have an ARM system and almost no TTSes work out of the box.
Some random thoughts: