4 papers · 1 filter
YingVideo-MV: Music-Driven Multi-Stage Video Generation
Jiahui Chen, Weida Wang, Runhua Shi +3
While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity c…
DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance
Junjie Zheng, Zihao Chen, Chaofan Ding +1
Current movie dubbing technology can generate the desired voice from a given speech prompt, ensuring good synchronization between speech and visuals while accurately conveying the…
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
Haomin Zhang, Chang Liu, Junjie Zheng +3
Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio ben…
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
Huan Yang, Jiahui Chen, Chaofan Ding +5
Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations th…