activity
20152023
most citedPhotorealistic Text-to-Image Diffusion Models with Deep Language Understanding

2.1k citations · 3.1k across the 18 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS20215 cited

WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis

Nanxin Chen, Yu Zhang, Heiga Zen +4

This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density…

eess.AS2021

Pushing the Limits of Non-Autoregressive Speech Recognition

Edwin G. Ng, Chung-Cheng Chiu, Yu Zhang +1

We combine recent advancements in end-to-end speech recognition to non-autoregressive automatic speech recognition. We push the limits of non-autoregressive state-of-the-art result…

eess.AS2020

WaveGrad: Estimating Gradients for Waveform Generation

Nanxin Chen, Yu Zhang, Heiga Zen +3

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and di…

eess.AS2020

Imputer: Sequence Modelling via Imputation and Dynamic Programming

William Chan, Chitwan Saharia, Geoffrey Hinton +2

This paper presents the Imputer, a neural sequence model that generates output sequences iteratively via imputations. The Imputer is an iterative generative model, requiring only a…

eess.AS20194 cited

SpecAugment on Large Scale Datasets

Daniel S. Park, Yu Zhang, Chung-Cheng Chiu +5

Recently, SpecAugment, an augmentation scheme for automatic speech recognition that acts directly on the spectrogram of input utterances, has shown to be highly effective in enhanc…

eess.AS20193 cited

Speaker-Targeted Audio-Visual Models for Speech Recognition in Cocktail-Party Environments

Guan-Lin Chao, William Chan, Ian Lane

Speech recognition in cocktail-party environments remains a significant challenge for state-of-the-art speech recognition systems, as it is extremely difficult to extract an acoust…