Skip to content

Selecting a Specific Speaker and Fine-Tuning for Multi-Speaker #51

Description

@sa-esf

Hi, and thank you for your amazing work on LLaSA TTS!

I have two questions regarding multi-speaker usage and training:

  1. I’m working with a multi-speaker version of the model, and I’d like to know how I can explicitly select a specific speaker at inference time. Is there a speaker_id parameter, speaker embedding, or reference audio I should provide to control the speaker output?

  2. I’m also preparing to fine-tune LLaSA for multiple speakers, and I’d like to know if there’s anything special I need to consider during data preparation or training. For example, should the dataset include a speaker_id field or follow a specific format to enable speaker conditioning?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions