9 papers
UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning
Hui Wang, Yifan Yang, Zeyue Tian +8
Audio generation and audio-to-text understanding remain largely separate, with diffusion models dominating high-fidelity synthesis and autoregressive (AR) language models driving c…
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
Hui Wang, Jinghua Zhao, Yifan Yang +9
Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scala…
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
Hui Wang, Jinghua Zhao, Junyang Cheng +5
Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only li…
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
Haoqin Sun, Chenyang Lyu, Xiangyu Kong +9
Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discr…
Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
Jinghua Zhao, Hang Su, Lichun Fan +4
With the rapid progress of large audio-language models (LALMs), audio question answering (AQA) has emerged as a challenging task requiring both fine-grained audio understanding and…
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
Haiying Xu, Haoze Liu, Mingshi Li +5
Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interact…