1 paper
Sirun Li, Minghao Liu, Ling Dai +4
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitra…