1 citations · 5 across the 20 of their papers we have counts for
Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
StepAudio 3 Realtime Technical Report
Bin Lin, Bo Zhao, Boyang Zhang +87
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a…
cs.SD2026
StepAudio 3 Gen Technical Report
Bin Lin, Bo Zhao, Boyang Wang +68
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe spee…
cs.SD2026
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
Qingjian Lin, Yuxin Li, Haoyang Zhang +14
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale languag…