4 papers · 1 filter
StepAudio 3 Realtime Technical Report
Bin Lin, Bo Zhao, Boyang Zhang +87
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a…
StepAudio 3 Gen Technical Report
Bin Lin, Bo Zhao, Boyang Wang +68
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe spee…
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
Yanru Wu, Jianning Wang, Chongxin Gan +1
Training general-purpose Audio Large Language Models (ALLMs) across diverse datasets is essential for holistic audio understanding, yet it faces significant challenges due to datas…
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
Sifei Li, Yang Li, Zizhou Wang +5
Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic…