58 citations · 74 across the 6 of their papers we have counts for
4 papers · 1 filter
StepAudio 3 Realtime Technical Report
Bin Lin, Bo Zhao, Boyang Zhang +87
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a…
StepAudio 3 Gen Technical Report
Bin Lin, Bo Zhao, Boyang Wang +68
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe spee…
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
Yong Ren, Jingbei Li, Haiyang Sun +6
Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence…
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
Fengrun Zhang, Wang Geng, Hukai Huang +3
In this paper, we introduce a speech-conditioned Large Language Model (LLM) integrated with a Mixture of Experts (MoE) based connector to address the challenge of Code-Switching (C…