RIDD introduces a novel plug-and-play framework for interpretable and responsible text-to-image (T2I) generation. Our approach enables fine-grained control over generated images by dynamically adjusting attributes such as age, gender, and race while maintaining fidelity and fairness in AI-generated content.
The rapid advancement in diffusion models has enabled high-quality text-to-image synthesis. However, interpretability, fairness, and responsible generation remain key challenges. RIDD addresses these challenges by:
- Providing fine-grained control over attributes like age, gender, and race.
- Using knowledge distillation and concept whitening for interpretable changes.
- Ensuring responsible AI by mitigating biases in generated images.
🔗 Live Demo: Coming Soon
🔗 Project Page: (Plug and Play Control)
If you use RIDD in your research, please cite:
@inproceedings{azam2025ridd,
author = {Basim Azam and Naveed Akhtar},
title = {Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2025},
url = {https://arxiv.org/abs/XXXX.XXXXX},
}We acknowledge contributions from
- The University of Melbourne
- OpenAI and Stable Diffusion
- Github Repos ()
