1 paper · 1 filter
Changye Li, Meng Lu, Yi Wu +1
While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and…