activity
20192022
most citedFastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

28 citations · 164 across the 30 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD20221 cited

3M: Multi-loss, Multi-path and Multi-level Neural Networks for speech recognition

Zhao You, Shulin Feng, Dan Su +1

Recently, Conformer based CTC/AED model has become a mainstream architecture for ASR. In this paper, based on our prior work, we identify and integrate several approaches to achiev…

cs.SD20211 cited

Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition

Max W. Y. Lam, Jun Wang, Chao Weng +2

End-to-end speech recognition generally uses hand-engineered acoustic features as input and excludes the feature extraction module from its joint optimization. To extract learnable…

cs.SD2021

SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts

Zhao You, Shulin Feng, Dan Su +1

Recently, Mixture of Experts (MoE) based Transformer has shown promising results in many domains. This is largely due to the following advantages of this architecture: firstly, MoE…

cs.SD2021

Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Rongzhi Gu, Shi-Xiong Zhang, Yuexian Zou +1

To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of the target speech in time-frequency domain under supervised…

cs.SD20211 cited

MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation

Xiyun Li, Yong Xu, Meng Yu +4

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance ove…

cs.SD2021

Generalized Spatio-Temporal RNN Beamformer for Target Speech Separation

Yong Xu, Zhuohuang Zhang, Meng Yu +2

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech i…