80 citations · 190 across the 38 of their papers we have counts for
15 papers · 1 filter
Audio-visual Multi-channel Recognition of Overlapped Speech
Jianwei Yu, Bo Wu, Rongzhi Gu +7
Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-…
Neural Spatio-Temporal Beamformer for Target Speech Separation
Yong Xu, Meng Yu, Shi-Xiong Zhang +4
Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmf…
Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning
Rongzhi Gu, Shi-Xiong Zhang, Lianwu Chen +5
Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. Howe…
Multi-modal Multi-channel Target Speech Separation
Rongzhi Gu, Shi-Xiong Zhang, Yong Xu +3
Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech sep…
Audio-visual Recognition of Overlapped speech for the LRS2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu +7
Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of…
Overlapped speech recognition from a jointly learned multi-channel neural speech extraction and representation
Bo Wu, Meng Yu, Lianwu Chen +3
We propose an end-to-end joint optimization framework of a multi-channel neural speech extraction and deep acoustic model without mel-filterbank (FBANK) extraction for overlapped s…