14 papers
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed +2
Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable…
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Sanghwan Kim, Rui Xiao, Stephan Alaniz +2
Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long v…
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
Ziqian Liu, Stephan Alaniz
Instruction-guided image editing has seen remarkable progress with models like FLUX.2 and Qwen-Image-Edit, yet they still struggle with complex scenarios with multiple similar inst…
Explaining CLIP Zero-shot Predictions Through Concepts
Onat Ozdemir, Anders Christensen, Stephan Alaniz +2
Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding.…
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
Leander Girrbach, Stephan Alaniz, Genevieve Smith +2
Vision-language models trained on large-scale multimodal datasets show strong demographic biases, but the role of training data in producing these biases remains unclear. A major b…
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
Rui Xiao, Sanghwan Kim, Yongqin Xian +2
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coa…