78 citations · 149 across the 11 of their papers we have counts for
1 paper · 2 filters
Ruihan Yang, Hannes Gamper, Sebastian Braun
We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the sy…