8 papers · 1 filter
Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection
Yunzhuo Chen, Jordan Vice, Naveed Akhtar +2
State-of-the-art text-to-image diffusion models can produce impressive visuals but may memorize and reproduce training images, creating copyright and privacy risks. Existing prompt…
CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
Li Liang, Bo Miao, Xinyu Wang +3
Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in th…
On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
Jordan Vice, Naveed Akhtar, Yansong Gao +2
Vision-Language Models (VLMs) are increasingly used as perceptual modules for visual content reasoning, including through captioning and DeepFake detection. In this work, we expose…
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
Jordan Vice, Naveed Akhtar, Leonid Sigal +2
The rapid proliferation of multimodal generative models has sparked critical discussions on their reliability, fairness and potential for misuse. While text-to-image models excel a…
Exploring Bias in over 100 Text-to-Image Generative Models
Jordan Vice, Naveed Akhtar, Richard Hartley +1
We investigate bias trends in text-to-image generative models over time, focusing on the increasing availability of models through open platforms like Hugging Face. While these pla…
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
Jordan Vice, Naveed Akhtar, Mubarak Shah +2
Training multimodal generative models on large, uncurated datasets can result in users being exposed to harmful, unsafe and controversial or culturally-inappropriate outputs. While…