1 paper · 1 filter
Ruihan Yang, Hannes Gamper, Sebastian Braun
We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the sy…