10 papers
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
Yaoteng Tan, Zikui Cai, M. Salman Asif
Controlling the behavior of text-to-image generative models is critical for safe and practical deployment. Existing safety approaches typically rely on model fine-tuning or curated…
Robust Multimodal Learning via Cross-Modal Proxy Tokens
Md Kaykobad Reza, Ameya Patil, Mashhour Solh +1
Multimodal models often experience a significant performance drop when one or more modalities are missing during inference. To address this challenge, we propose a simple yet effec…
Cross-Modal Safety Alignment: Is textual unlearning all you need?
Trishna Chakraborty, Erfan Shayegani, Zikui Cai +5
Recent studies reveal that integrating new modalities into Large Language Models (LLMs), such as Vision-Language Models (VLMs), creates a new attack surface that bypasses existing…
EigenScore: OOD Detection using Covariance in Diffusion Models
Shirin Shoushtari, Yi Wang, Xiao Shi +2
Out-of-distribution (OOD) detection is critical for the safe deployment of machine learning systems in safety-sensitive domains. Diffusion models have recently emerged as powerful…
Gaussian is All You Need: A Unified Framework for Solving Inverse Problems via Diffusion Posterior Sampling
Nebiyou Yismaw, Ulugbek S. Kamilov, M. Salman Asif
Diffusion models can generate a variety of high-quality images by modeling complex data distributions. Trained diffusion models can also be very effective image priors for solving…
VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta +5
Human pose and shape (HPS) estimation methods have been extensively studied, with many demonstrating high zero-shot performance on in-the-wild images and videos. However, these met…