collaborators

6 papers

cs.SD2025

A Neural Speech Codec for Noise Robust Speech Coding

Jiayi Huang, Zeyu Yan, Wenbin Jiang +2

This paper considers the joint compression and enhancement problem for speech signal in the presence of noise. Recently, the SoundStream codec, which relies on end-to-end joint tra…

eess.AS2025

ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark

He Wang, Linhan Ma, Dake Guo +4

Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…

cs.SD2025

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding

Dake Guo, Jixun Yao, Linhan Ma +2

Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interacti…

eess.AS2025

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech

Linhan Ma, Dake Guo, He Wang +2

Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in…

cs.SD2025

Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization

Weifei Jin, Junjie Su, Hejia Wang +2

With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adver…

cs.SD2025

The ICME 2025 Audio Encoder Capability Challenge

Junbo Zhang, Heinrich Dinkel, Qiong Song +8

This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit…