1 paper
Yufei He, Yuan Sui, Xiaoxin He +3
Existing foundation models, such as CLIP, aim to learn a unified embedding space for multimodal data, enabling a wide range of downstream web-based applications like search, recomm…