activity
20182024
most citedGenerative adversarial network-based glottal waveform model for statistical parametric speech synthesis

50 citations · 52 across the 4 of their papers we have counts for

collaborators

12 papers

cs.SD2024

Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis

Lauri Juvela, Xin Wang

Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to t…

eess.AS2020

Conditional Spoken Digit Generation with StyleGAN

Kasperi Palkama, Lauri Juvela, Alexander Ilin

This paper adapts a StyleGAN model for speech generation with minimal or no conditioning on text. StyleGAN is a multi-scale convolutional GAN capable of hierarchically capturing da…

eess.AS20191 cited

Transferring neural speech waveform synthesizers to musical instrument sounds generation

Yi Zhao, Xin Wang, Lauri Juvela +1

Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different meth…

eess.AS2019

ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech

Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37

Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…

eess.AS20191 cited

GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram

Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi +1

Recent advances in neural network -based text-to-speech have reached human level naturalness in synthetic speech. The present sequence-to-sequence models can directly map text to m…

eess.AS201950 cited

Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis

Bajibabu Bollepalli, Lauri Juvela, Paavo Alku

Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glot…