activity
20182022
most citedEnd-to-End Multi-Channel Speech Separation

80 citations · 180 across the 15 of their papers we have counts for

collaborators
Showing eess.ASShow all

9 papers · 1 filter

eess.AS202228 cited

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Rongjie Huang, Max W. Y. Lam, Jun Wang +4

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…

eess.AS202226 cited

BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

Max W. Y. Lam, Jun Wang, Dan Su +1

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral d…

eess.AS202219 cited

DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Songxiang Liu, Dan Su, Dong Yu

Denoising diffusion probabilistic models (DDPMs) are expressive generative models that have been used to solve a variety of speech synthesis problems. However, because of their hig…

eess.AS2021

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

Songxiang Liu, Shan Yang, Dan Su +1

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSS…

eess.AS20213 cited

Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis

Jian Cong, Shan Yang, Lei Xie +1

Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectr…

eess.AS2021

Controllable Context-aware Conversational Speech Synthesis

Jian Cong, Shan Yang, Na Hu +3

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interloc…