Skip to content

Great comparison #1

Description

@jerkstorecaller

Thanks for making this, it's so nice to be able to sample all the TTSes in one place without having to get the models. It's appreciated. I have an ARM system and almost no TTSes work out of the box.

Some random thoughts:

  • A new TTS just came out, dots.tts 2B. If the demos aren't cherry-picked this is the best TTS I've ever heard, including best voice cloning and best expressivity. You really want to add this. https://rednote-hilab.github.io/dots.tts-demo/
  • Could you add MOSS TTS 1.5 which came out last month? It's not as impressive but it's pretty solid.
  • Why are there only partial entries on some prompts? For example prompt 5 has way fewer entries than prompt 3. Does it take a long time to do a full run with all supported models?
  • Is there some source that tracks WER for all models on the same benchmark? And if so, would you want to add it in a column on the Listen page, next to the Play Audio button for each model? This way, it helps quickly see the trade-offs of a great-sounding model. For example your 'TLDR best cloning choice,' OmniVoice, on Prompt 3 voice clone skips the entire beginning of the sentence (5 words). It feels wrong to make this your Great comparison #1 recommendation when it has such a glaring flaw, I don't care how good the voice is.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions