1 paper
Feixiang Liu, Likun Wang, Qiang Qiu +3
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the…