Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
Bingli Wang, Huanze Tang, Haijun Lv +5
In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmar…
cs.CV2025
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
Zhifang Zhang, Qiqi Tao, Jiaqi Lv +3
Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing b…
cs.CV2025
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
Zongyu Wu, Yuwei Niu, Hongcheng Gao +12
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. E…