1 paper
Guangyu Sun, Shlok Kumar Mishra, Wentao Bao +8
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream ge…