6 papers
A Neural Speech Codec for Noise Robust Speech Coding
Jiayi Huang, Zeyu Yan, Wenbin Jiang +2
This paper considers the joint compression and enhancement problem for speech signal in the presence of noise. Recently, the SoundStream codec, which relies on end-to-end joint tra…
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
He Wang, Linhan Ma, Dake Guo +4
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
Dake Guo, Jixun Yao, Linhan Ma +2
Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interacti…
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
Linhan Ma, Dake Guo, He Wang +2
Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in…
Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization
Weifei Jin, Junjie Su, Hejia Wang +2
With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adver…
The ICME 2025 Audio Encoder Capability Challenge
Junbo Zhang, Heinrich Dinkel, Qiong Song +8
This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit…