1 paper
Wenyu Zhang, Wei En Ng, Lixin Ma +5
Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessa…