5 papers
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
Weihao Wu, Liang Cao, Xinyu Wu +4
Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create im…
LeVo: High-Quality Song Generation with Multi-Preference Alignment
Shun Lei, Yaoxun Xu, Zhiwei Lin +10
Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing…
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Weihao wu, Zhiwei Lin, Yixuan Zhou +6
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
Wei Chen, Xintao Zhao, Jun Chen +3
Singing voice conversion (SVC) is hindered by noise sensitivity due to the use of non-robust methods for extracting pitch and energy during the inference. As clean signals are key…
MuCodec: Ultra Low-Bitrate Music Codec
Yaoxun Xu, Hangting Chen, Jianwei Yu +5
Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity…