activity
20192022
most citedFastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

28 citations · 81 across the 9 of their papers we have counts for

collaborators

9 papers

eess.AS202228 cited

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Rongjie Huang, Max W. Y. Lam, Jun Wang +4

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…

eess.AS202226 cited

BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

Max W. Y. Lam, Jun Wang, Dan Su +1

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral d…

cs.LG202111 cited

Bilateral Denoising Diffusion Models

Max W. Y. Lam, Jun Wang, Rongjie Huang +2

Denoising diffusion probabilistic models (DDPMs) have emerged as competitive generative models yet brought challenges to efficient sampling. In this paper, we propose novel bilater…

cs.SD20211 cited

Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition

Max W. Y. Lam, Jun Wang, Chao Weng +2

End-to-end speech recognition generally uses hand-engineered acoustic features as input and excludes the feature extraction module from its joint optimization. To extract learnable…

eess.AS20214 cited

Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation

Max W. Y. Lam, Jun Wang, Dan Su +1

One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throu…

eess.AS2021

Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect

Jun Wang, Max W. Y. Lam, Dan Su +1

We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns…