Showing cs.SDShow all
3 papers · 1 filter
cs.SD2025
Index-ASR Technical Report
Zheshu Song, Lu Wang, Wei Deng +3
Automatic speech recognition (ASR) has witnessed remarkable progress in recent years, largely driven by the emergence of LLM-based ASR paradigm. Despite their strong performance on…
cs.SD2025
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…
cs.SD2024
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
Ziya Zhou, Yuhang Wu, Zhiyue Wu +7
Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the s…