1 paper
Kaito Watanabe, Taisei Yamamoto, Tomoki Doi +1
One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we…