1 paper
Chanyoung Gwak, Yoonwoo Jeong, Byungwoo Jeon +3
Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly…