collaborators

8 papers

cs.SD2026

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

Qingjian Lin, Yuxin Li, Haoyang Zhang +14

Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale languag…

cs.CL2026

Chronological Thinking in Full-Duplex Spoken Dialogue Language Models

Donghang Wu, Haoyang Zhang, Chen Chen +8

Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user…

eess.AS2026

Step-Audio-R1.5 Technical Report

Yuxin Zhang, Xiangyu Tony Zhang, Daijiao Liu +16

Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic…

cs.CL2026

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Donghang Wu, Haoyang Zhang, Jun Chen +9

Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entire thought process sequentially.…

cs.AI2025

Step-Audio-R1 Technical Report

Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…

cs.CL2025

Step-Audio-EditX Technical Report

Chao Yan, Boyong Wu, Peng Yang +12

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguisti…