1 paper
Bridget Leonard, Scott O. Murray
Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These erro…