1 paper
Yuan Zhao, Zhenqi Jia, Rui Liu +3
Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual inf…