106 citations · 165 across the 16 of their papers we have counts for
7 papers · 1 filter
MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
Xiyun Li, Yong Xu, Meng Yu +4
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance ove…
Speaker and Direction Inferred Dual-channel Speech Separation
Chenxing Li, Jiaming Xu, Nima Mesgarani +1
Most speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough general- ization capabilities for real scenarios where the n…
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou +1
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and perfor…
Audio-visual Speech Separation with Adversarially Disentangled Visual Representation
Peng Zhang, Jiaming Xu, Jing shi +2
Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they bu…
Unsupervised pre-training for sequence to sequence speech recognition
Zhiyun Fan, Shiyu Zhou, Bo Xu
This paper proposes a novel approach to pre-train encoder-decoder sequence-to-sequence (seq2seq) model with unpaired speech and transcripts respectively. Our pre-training method is…
Single-channel Speech Dereverberation via Generative Adversarial Training
Chenxing Li, Tieqiang Wang, Shuang Xu +1
In this paper, we propose a single-channel speech dereverberation system (DeReGAT) based on convolutional, bidirectional long short-term memory and deep feed-forward neural network…