16 papers
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
Haomin Zhang, Kristin Qi, Shuxin Yang +3
Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned…
Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
Xinhan Di, JoyJiaoW
Reinforcement learning scaling enhances the reasoning capabilities of large language models, with reinforcement learning serving as the key technique to draw out complex reasoning.…
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
Wenjie Tian, Xinfa Zhu, Haohe Liu +6
While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This…
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
Chang Liu, Haomin Zhang, Shiyu Xia +5
Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, e…
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Junjie Zheng, Zihao Chen, Chaofan Ding +5
Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conv…
Towards Film-Making Production Dialogue, Narration, Monologue Adaptive Moving Dubbing Benchmarks
Chaoyi Wang, Junjie Zheng, Zihao Chen +6
Movie dubbing has advanced significantly, yet assessing the real-world effectiveness of these models remains challenging. A comprehensive evaluation benchmark is crucial for two ke…