1 citations · 1 across the 3 of their papers we have counts for
4 papers
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
Qingjian Lin, Yuxin Li, Haoyang Zhang +14
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale languag…
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
Haoyang Zhang, Jun Chen, Donghang Wu +13
Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses.…
Step-Audio-R1 Technical Report
Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14
Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…
Step-Audio 2 Technical Report
Boyong Wu, Chao Yan, Chen Hu +106
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…