3 papers
cs.SD2025
Index-ASR Technical Report
Zheshu Song, Lu Wang, Wei Deng +3
Automatic speech recognition (ASR) has witnessed remarkable progress in recent years, largely driven by the emergence of LLM-based ASR paradigm. Despite their strong performance on…
eess.AS2025
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
Jinming Chen, Lu Wang, Zheshu Song +1
Driven by large scale datasets and LLM based architectures, automatic speech recognition (ASR) systems have achieved remarkable improvements in accuracy. However, challenges persis…
cs.SD2025
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…