activity
20192022
most citedFastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

28 citations · 81 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS202228 cited

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Rongjie Huang, Max W. Y. Lam, Jun Wang +4

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…

eess.AS202226 cited

BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

Max W. Y. Lam, Jun Wang, Dan Su +1

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral d…

eess.AS20214 cited

Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation

Max W. Y. Lam, Jun Wang, Dan Su +1

One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throu…

eess.AS2021

Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect

Jun Wang, Max W. Y. Lam, Dan Su +1

We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns…

eess.AS2021

Contrastive Separative Coding for Self-supervised Representation Learning

Jun Wang, Max W. Y. Lam, Dan Su +1

To extract robust deep representations from long sequential modeling of speech data, we propose a self-supervised learning approach, namely Contrastive Separative Coding (CSC). Our…

eess.AS20219 cited

Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks

Max W. Y. Lam, Jun Wang, Dan Su +1

Recent research on the time-domain audio separation networks (TasNets) has brought great success to speech separation. Nevertheless, conventional TasNets struggle to satisfy the me…