1 paper
Ruoxi Cheng, Yizhong Ding, Jian Zhao +5
Contrastive pretraining models such as CLIP and CLAP, serve as the ubiquitous perceptual backbones for modern multimodal large models, yet their reliance on web-scale data raises g…