2 papers
eess.AS2025
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
Jun Wang, Xijuan Zeng, Chunyu Qiang +19
We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce m…
cs.SD2024
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
Wangjin Zhou, Fengrun Zhang, Yiming Liu +3
This study presents an innovative Zero-Shot any-to-any Singing Voice Conversion (SVC) method, leveraging a novel clustering-based phoneme representation to effectively separate con…