2 papers
eess.AS2022
TransFusion: Transcribing Speech with Multinomial Diffusion
Matthew Baas, Kevin Eloff, Herman Kamper
Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional t…
cs.SD2022
GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from Diffusion Models
Matthew Baas, Herman Kamper
We propose AudioStyleGAN (ASGAN), a new generative adversarial network (GAN) for unconditional speech synthesis. As in the StyleGAN family of image synthesis models, ASGAN maps sam…