Showing cs.SDShow all
3 papers · 1 filter
cs.SD2025
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…
cs.SD2025
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
Haoran Zhou, Xingchen Song, Brendan Fahy +9
OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence…
cs.SD2025
MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
Hao Zhou, Xiaobao Guo, Yuzhe Zhu +1
Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual…