Skip to content

Explanation on ControlNet model structure #2

Description

@JamesLong199

Thank you for your awesome work, I would like to ask a question about the model structure:

there is a VAE encoder for the ControlNetXS model, which compress and samples the camera condition. in addition, this VAE encoder is trained on diffusion loss only, without KL loss. I was wondering what is the motivation behind using this VAE encoder?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions