50 citations · 52 across the 4 of their papers we have counts for
12 papers
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
Lauri Juvela, Xin Wang
Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to t…
Conditional Spoken Digit Generation with StyleGAN
Kasperi Palkama, Lauri Juvela, Alexander Ilin
This paper adapts a StyleGAN model for speech generation with minimal or no conditioning on text. StyleGAN is a multi-scale convolutional GAN capable of hierarchically capturing da…
Transferring neural speech waveform synthesizers to musical instrument sounds generation
Yi Zhao, Xin Wang, Lauri Juvela +1
Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different meth…
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…
GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram
Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi +1
Recent advances in neural network -based text-to-speech have reached human level naturalness in synthetic speech. The present sequence-to-sequence models can directly map text to m…
Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis
Bajibabu Bollepalli, Lauri Juvela, Paavo Alku
Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glot…