1 paper · 1 filter
Sirun Li, Minghao Liu, Ling Dai +4
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitra…