activity
20182022
most citedTraditional Machine Learning for Pitch Detection

32 citations · 68 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

10 papers · 1 filter

eess.AS2022

Distribution augmentation for low-resource expressive text-to-speech

Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…

eess.AS2021

Multi-Scale Spectrogram Modelling for Neural Text-to-Speech

Ammar Abbas, Bajibabu Bollepalli, Alexis Moinet +6

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrog…

eess.AS2021

A learned conditional prior for the VAE acoustic space of a TTS system

Penny Karanasou, Sri Karlapati, Alexis Moinet +5

Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…

eess.AS20204 cited

Parallel WaveNet conditioned on VAE latent vectors

Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…

eess.AS2020

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

Sri Karlapati, Ammar Abbas, Zack Hodari +4

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn…

eess.AS2020

CAMP: a Two-Stage Approach to Modelling Prosody in Context

Zack Hodari, Alexis Moinet, Sri Karlapati +6

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody…