1 paper
Yeming Chen, Siyu Zhang, Yaoru Sun +2
With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…