most citedsVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS2024

SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech

Jingru Lin, Meng Ge, Junyi Ao +2

It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on singl…

cs.SD20241 cited

sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

Qu Yang, Qianhui Liu, Nan Li +3

Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neu…

eess.AS20241 cited

An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement

Qiquan Zhang, Meng Ge, Hongxu Zhu +4

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to…

eess.AS2023

The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge

Meng Ge, Yizhou Peng, Yidi Jiang +6

This paper summarizes our team's efforts in both tracks of the ICMC-ASR Challenge for in-car multi-channel automatic speech recognition. Our submitted systems for ICMC-ASR Challeng…

eess.AS2023

Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech

Jingru Lin, Meng Ge, Wupeng Wang +2

Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-label…