4 papers
Masked Audio Modeling with CLAP and Multi-Objective Learning
Yifei Xin, Xiulian Peng, Yan Lu
Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accou…
Improving Text-Audio Retrieval by Text-aware Attention Pooling and Prior Matrix Revised Loss
Yifei Xin, Dongchao Yang, Yuexian Zou
In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames w…
Improving Weakly Supervised Sound Event Detection with Causal Intervention
Yifei Xin, Dongchao Yang, Fan Cui +2
Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurr…
Improving Speech Enhancement via Event-based Query
Yifei Xin, Xiulian Peng, Yan Lu
Existing deep learning based speech enhancement (SE) methods either use blind end-to-end training or explicitly incorporate speaker embedding or phonetic information into the SE ne…