38 papers
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
Sarah Rastegar, Mina Ghadimi Atigh, Pascal Mettes +2
Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typically based on self-supervision and cont…
See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs
Yuqing Lei, Wenbo Lyu, Yingjun Du +3
Large Vision-Language Models (LVLMs) excel at multimodal tasks but remain prone to object hallucinations. Prior training-free remedies often uniformly strengthen visual signals, wh…
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Zhen Zeng, Leijiang Gu, Zhangling Duan +4
Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearning methods can remove such content, they o…
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
Ziqi Wang, Chang Che, Qi Wang +4
While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safe…
Elastic ViTs from Pretrained Models without Retraining
Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort +2
Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world con…
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
RÄzvan-Andrei MatiÅan, Vincent Tao Hu, Grigory Bartosh +6
We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous tra…