1 paper · 1 filter
Haochen You, Baojing Liu
Recent advances in multimodal learning have largely relied on pairwise contrastive objectives to align different modalities, such as text, video, and audio, in a shared embedding s…