2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Ryan Spencer, Roey Yaari, Ritvik Vemavarapu +3
Multimodal large language models (MLLMs) are proficient in perception and instruction-following, but they still struggle with spatial reasoning: the ability to mentally track and m…