1 paper · 1 filter
Feixiang Liu, Likun Wang, Qiang Qiu +3
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the…