1 paper · 1 filter
Chuang Ma, Qianying Liu, Tomoyuki Obuchi +6
Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly attended visual information. In t…