4 papers
Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
Zhongweiyang Xu, Ashutosh Pandey, Juan Azcarreta +4
We propose Uni-ArrayDPS, a novel diffusion-based refinement framework for unified multi-channel speech enhancement and separation. Existing methods for multi-channel speech enhance…
ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
Zhongweiyang Xu, Ashutosh Pandey, Juan Azcarreta +3
Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-…
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
Jiahui Zhao, Hao Shi, Chenrui Cui +5
Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. A…
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
Hao Shi, Yuan Gao, Zhaoheng Ni +1
Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy…