1 paper · 1 filter
Bob Zhang, Haoran Li, Tao Zhang +5
Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instruct…