4 papers
Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss
Nian Shao, Xian Li, Xiaofei Li
Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled data. Recent systems leverage large pretrai…
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
Nian Shao, Rui Zhou, Pengyu Wang +4
In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) p…
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
Rui Zhou, Xian Li, Ying Fang +1
In this work, we propose Mel-FullSubNet, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (…
Fine-tune the pretrained ATST model for sound event detection
Nian Shao, Xian Li, Xiaofei Li
Sound event detection (SED) often suffers from the data deficiency problem. The recent baseline system in the DCASE2023 challenge task 4 leverages the large pretrained self-supervi…