I installed BreezyVoice locally on a PC without a GPU, but the voice cloning results are noticeably different from the Hugging Face version. The local output has unnatural pronunciation, often fails to read the text properly, and sometimes blends the sample voice directly into the synthesized output.
The reference voice sample is 6–10 seconds long, and the input text for synthesis is around 300 characters.
Also, how can I set a random seed like the option available on the Hugging Face demo?
I installed BreezyVoice locally on a PC without a GPU, but the voice cloning results are noticeably different from the Hugging Face version. The local output has unnatural pronunciation, often fails to read the text properly, and sometimes blends the sample voice directly into the synthesized output.
The reference voice sample is 6–10 seconds long, and the input text for synthesis is around 300 characters.
Also, how can I set a random seed like the option available on the Hugging Face demo?