13 papers
Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
Donghang Wu, Haoyang Zhang, Jun Chen +9
Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entire thought process sequentially.…
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
Hexin Liu, Haoyang Zhang, Qiquan Zhang +4
Code-switching automatic speech recognition (CS-ASR) presents unique challenges due to language confusion introduced by spontaneous intra-sentence switching and accent bias that bl…
Step-Audio-R1 Technical Report
Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14
Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…
Step-Audio-EditX Technical Report
Chao Yan, Boyong Wu, Peng Yang +12
We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguisti…
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
Hexin Liu, Xiangyu Zhang, Haoyang Zhang +4
Code-switching (CS) refers to the switching of languages within a speech signal and results in language confusion for automatic speech recognition (ASR). To address language confus…
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
Yayue Deng, Guoqiang Hu, Haiyang Sun +6
Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on…