9 papers
On the Emotion Understanding of Synthesized Speech
Yuan Ge, Haishu Zhao, Aokai Hao +10
Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesize…
StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control
Haishu Zhao, Aokai Hao, Yuan Ge +3
Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For mor…
M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
Ruixiang Mao, Xiangnan Ma, Qing Yang +7
The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mappin…
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
Jianjin Wang, Runsong Zhao, Xiaoqian Liu +6
Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so…
FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
Yuan Ge, Saihan Chen, Jingqi Xiao +5
Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking…
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
Yanfan Du, Jun Zhang, Bin Wang +6
Recent advances in speech large language models (SLMs) have improved speech recognition and translation in general domains, but accurately generating domain-specific terms or neolo…