6 papers
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
Step-Audio-EditX Technical Report
Chao Yan, Boyong Wu, Peng Yang +12
We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguisti…
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
Yayue Deng, Guoqiang Hu, Haiyang Sun +6
Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on…
Step-Audio 2 Technical Report
Boyong Wu, Chao Yan, Chen Hu +106
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…
Recent Advances in End-to-End Simultaneous Speech Translation
Xiaoqian Liu, Guoqiang Hu, Yangfan Du +5
Simultaneous speech translation (SimulST) is a demanding task that involves generating translations in real-time while continuously processing speech input. This paper offers a com…
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
Haoyu Wang, Guoqiang Hu, Guodong Lin +2
As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its…