7 papers
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
Wenjie Tian, Bingshen Mu, Guobin Ma +3
Automatic speech recognition (ASR) systems based on large language models (LLMs) achieve superior performance by leveraging pretrained LLMs as decoders, but their token-by-token ge…
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
Guobin Ma, Yuxuan Xia, Jixun Yao +5
This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The…
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
Junjie Zheng, Chunbo Hao, Guobin Ma +5
Singing Voice Synthesis (SVS) remains constrained in practical deployment due to its strong dependence on accurate phoneme-level alignment and manually annotated melody contours, r…
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
Zhao Guo, Ziqian Ning, Guobin Ma +1
Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming…
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
Guobin Ma, Jixun Yao, Ziqian Ning +4
Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand…
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
Xuelong Geng, Qijie Shao, Hongfei Xue +20
Empathy is crucial in enabling natural interactions within spoken dialogue systems, allowing machines to recognize and respond appropriately to paralinguistic cues such as age, gen…