1 paper
Xinhan Di, Kristin Qi, Pengqian Yu
Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly gen…