1 paper · 1 filter
Zhipeng Cai, Zhuang Liu, Yunyang Xiong +3
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D u…