2 citations · 2 across the 3 of their papers we have counts for
4 papers
StepAudio 3 Realtime Technical Report
Bin Lin, Bo Zhao, Boyang Zhang +87
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
Ziya Zhou, Yuhang Wu, Zhiyue Wu +7
Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the s…