4 papers
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
Improved Training Technique for Latent Consistency Models
Quan Dao, Khanh Doan, Di Liu +2
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demo…
Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation
Anh Bui, Long Vuong, Khanh Doan +4
Diffusion models excel at generating visually striking content from text but can inadvertently produce undesirable or harmful content when trained on unfiltered internet data. A pr…
Connective Viewpoints of Signal-to-Noise Diffusion Models
Khanh Doan, Long Tung Vuong, Tuan Nguyen +5
Diffusion models (DM) have become fundamental components of generative models, excelling across various domains such as image creation, audio generation, and complex data interpola…