collaborators

5 papers

cs.CL2025

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

Weihao Wu, Liang Cao, Xinyu Wu +4

Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create im…

cs.SD2025

LeVo: High-Quality Song Generation with Multi-Preference Alignment

Shun Lei, Yaoxun Xu, Zhiwei Lin +10

Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing…

cs.SD2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

Weihao wu, Zhiwei Lin, Yixuan Zhou +6

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…

cs.SD2024

RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion

Wei Chen, Xintao Zhao, Jun Chen +3

Singing voice conversion (SVC) is hindered by noise sensitivity due to the use of non-robust methods for extracting pitch and energy during the inference. As clean signals are key…

cs.SD2024

MuCodec: Ultra Low-Bitrate Music Codec

Yaoxun Xu, Hangting Chen, Jianwei Yu +5

Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity…