1 citations · 2 across the 4 of their papers we have counts for
5 papers
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
Jingru Lin, Meng Ge, Junyi Ao +2
It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on singl…
sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
Qu Yang, Qianhui Liu, Nan Li +3
Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neu…
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
Qiquan Zhang, Meng Ge, Hongxu Zhu +4
Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to…
The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge
Meng Ge, Yizhou Peng, Yidi Jiang +6
This paper summarizes our team's efforts in both tracks of the ICMC-ASR Challenge for in-car multi-channel automatic speech recognition. Our submitted systems for ICMC-ASR Challeng…
Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
Jingru Lin, Meng Ge, Wupeng Wang +2
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-label…