31 papers
DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation
Héctor Laria, Yiping Han, Julian D. Santamaria +4
Adapting pre-trained text-to-image diffusion models, whether to learn new visual concepts or erase unwanted ones, is routinely evaluated on its intended effects alone. We argue thi…
Position: Modular Memory is the Key to Continual Learning Agents
Vaggelis Dorovatas, Malte Schwerin, Andrew D. Bagdanov +21
Foundation models have transformed machine learning through large-scale pretraining and increased test-time compute. Despite surpassing human performance in several domains, these…
Group Preference Collapse in Personalized Multimodal Large Language Models
Fan Lyu, Wenqi Zhang, Joost van de Weijer
Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse us…
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
Simone Magistri, Dipam Goswami, Marco Mistretta +3
Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individual modality encoders are applie…
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
Yuyang Liu, Qiuhe Hong, Linlan Huang +6
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerfu…
Adversarial Concept Distillation for One-Step Diffusion Personalization
Yixiong Yang, Tao Wu, Senmao Li +4
Recent progress in accelerating text-to-image diffusion models enables high-fidelity synthesis within a single denoising step. However, customizing the fast one-step models remains…