1 paper · 1 filter
Saurav Jha, M. Jehanzeb Mirza, Wei Lin +2
Vision-Language Models (VLMs) remain limited in spatial reasoning tasks that require multi-view understanding and embodied perspective shifts. Recent approaches such as MindJourney…