1 paper
Haomin Zhang, Sizhe Shan, Haoyu Wang +4
Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-s…