1 paper
Yilin Gao, Kangyi Chen, Zhongxing Peng +2
Current visual foundation models (VFMs) face a fundamental limitation in transferring knowledge from vision language models (VLMs), while VLMs excel at modeling cross-modal interac…