Showing cs.MMShow all
2 papers · 1 filter
cs.MM2025
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
Wenjie Tian, Xinfa Zhu, Haohe Liu +6
While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This…
cs.MM2025
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Junjie Zheng, Zihao Chen, Chaofan Ding +5
Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conv…