2 papers
cs.SD2025
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
Chang Liu, Haomin Zhang, Shiyu Xia +5
Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, e…
cs.CV2025
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
Haomin Zhang, Chang Liu, Junjie Zheng +3
Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio ben…