1 paper
Chao Zhou, Yiling Chen, Qi Chu +3
Although pretrained joint audio-visual diffusion models offer rich control over \emph{what} to generate, they provide no explicit control over \emph{when} an utterance should occur…