1 paper · 1 filter
Suchae Jeong, Jaehwi Song, Haeone Lee +9
Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems…