94 citations · 119 across the 8 of their papers we have counts for
7 papers · 1 filter
Bilateral Denoising Diffusion Models
Max W. Y. Lam, Jun Wang, Rongjie Huang +2
Denoising diffusion probabilistic models (DDPMs) have emerged as competitive generative models yet brought challenges to efficient sampling. In this paper, we propose novel bilater…
Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition
Max W. Y. Lam, Jun Wang, Chao Weng +2
End-to-end speech recognition generally uses hand-engineered acoustic features as input and excludes the feature extraction module from its joint optimization. To extract learnable…
PanGu-: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
Wei Zeng, Xiaozhe Ren, Teng Su +35
Large-scale Pretrained Language Models (PLMs) have become the new paradigm for Natural Language Processing (NLP). PLMs with hundreds of billions parameters such as GPT-3 have demon…
Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation
Max W. Y. Lam, Jun Wang, Dan Su +1
One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throu…
Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect
Jun Wang, Max W. Y. Lam, Dan Su +1
We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns…
Contrastive Separative Coding for Self-supervised Representation Learning
Jun Wang, Max W. Y. Lam, Dan Su +1
To extract robust deep representations from long sequential modeling of speech data, we propose a self-supervised learning approach, namely Contrastive Separative Coding (CSC). Our…