1 paper · 1 filter
Yipu Wang, Yuheng Ji, Yuyang Liu +10
Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), espe…