Thank you for your awesome work, I would like to ask a question about the model structure:
there is a VAE encoder for the ControlNetXS model, which compress and samples the camera condition. in addition, this VAE encoder is trained on diffusion loss only, without KL loss. I was wondering what is the motivation behind using this VAE encoder?
Thank you for your awesome work, I would like to ask a question about the model structure:
there is a VAE encoder for the ControlNetXS model, which compress and samples the camera condition. in addition, this VAE encoder is trained on diffusion loss only, without KL loss. I was wondering what is the motivation behind using this VAE encoder?