1 paper · 1 filter
Yitong Chen, Lingchen Meng, Wujian Peng +4
Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowi…