5 papers
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang, Ziyang Ma, Yizhou Peng +26
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
Yongqi Wang, Ruofan Hu, Rongjie Huang +6
Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized…
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
Yongqi Wang, Wenxiang Guo, Rongjie Huang +5
Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency…
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
Fuming You, Minghui Fang, Li Tang +3
Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motio…
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
Zhiqing Hong, Rongjie Huang, Xize Cheng +5
A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid t…