18 citations · 39 across the 11 of their papers we have counts for
1 paper · 1 filter
Patrick Rim, Tom Long, Ekta Prashnani +6
Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as…