1 paper
Svetlana Kutuzova, Oswin Krause, Douglas McCloskey +2
Multimodal generative models should be able to learn a meaningful latent representation that enables a coherent joint generation of all modalities (e.g., images and text). Many app…