106 citations · 165 across the 16 of their papers we have counts for
26 papers
WASE: Learning When to Attend for Speaker Extraction in Cocktail Party Environments
Yunzhe Hao, Jiaming Xu, Peng Zhang +1
In the speaker extraction problem, it is found that additional information from the target speaker contributes to the tracking and extraction of the target speaker, which includes…
MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
Xiyun Li, Yong Xu, Meng Yu +4
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance ove…
Speaker and Direction Inferred Dual-channel Speech Separation
Chenxing Li, Jiaming Xu, Nima Mesgarani +1
Most speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough general- ization capabilities for real scenarios where the n…
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou +1
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and perfor…
CIF-based Collaborative Decoding for End-to-end Contextual Speech Recognition
Minglun Han, Linhao Dong, Shiyu Zhou +1
End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure…
Audio-visual Speech Separation with Adversarially Disentangled Visual Representation
Peng Zhang, Jiaming Xu, Jing shi +2
Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they bu…