4 papers
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation
Anh Bui, Long Vuong, Khanh Doan +4
Diffusion models excel at generating visually striking content from text but can inadvertently produce undesirable or harmful content when trained on unfiltered internet data. A pr…
Improved Training Technique for Latent Consistency Models
Quan Dao, Khanh Doan, Di Liu +2
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demo…
Hiding and Recovering Knowledge in Text-to-Image Diffusion Models via Learnable Prompts
Anh Bui, Khanh Doan, Trung Le +3
Diffusion models have demonstrated remarkable capability in generating high-quality visual content from textual descriptions. However, since these models are trained on large-scale…