1 paper
Ryan Spencer, Roey Yaari, Ritvik Vemavarapu +3
Multimodal large language models (MLLMs) are proficient in perception and instruction-following, but they still struggle with spatial reasoning: the ability to mentally track and m…