1 paper
Siquan Huang, Yijiang Li, Ningzhi Gao +3
Self-supervised and multimodal vision encoders learn strong visual representations that are widely adopted in downstream vision tasks and large vision-language models (LVLMs). Howe…