1 paper
Xingyu Zhu, Liang Yi, Shuo Wang +4
Multimodal 3D vision-language models show strong generalization across diverse 3D tasks, but their performance still degrades notably under domain shifts. This has motivated recent…