4 papers
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Junjie Zheng, Zihao Chen, Chaofan Ding +5
Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conv…
DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos
Yunming Liang, Zihao Chen, Chaofan Ding +1
Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment b…
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
Zihao Chen, Haomin Zhang, Xinhan Di +10
Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-…
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
Xinhan Di, Zihao Chen, Yunming Liang +3
Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we pro…