activity
20182022
most citedTraditional Machine Learning for Pitch Detection

32 citations · 68 across the 7 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

eess.AS20204 cited

Parallel WaveNet conditioned on VAE latent vectors

Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…

eess.AS2020

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

Sri Karlapati, Ammar Abbas, Zack Hodari +4

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn…

eess.AS2020

CAMP: a Two-Stage Approach to Modelling Prosody in Context

Zack Hodari, Alexis Moinet, Sri Karlapati +6

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody…

eess.AS20201 cited

Glottal source estimation robustness: A comparison of sensitivity of voice source estimation techniques

Thomas Drugman, Thomas Dubuisson, Alexis Moinet +2

This paper addresses the problem of estimating the voice source directly from speech waveforms. A novel principle based on Anticausality Dominated Regions (ACDR) is used to estimat…

eess.AS2020

CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech

Sri Karlapati, Alexis Moinet, Arnaud Joly +3

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects l…

cs.SD202031 cited

Voice Conversion for Whispered Speech Synthesis

Marius Cotescu, Thomas Drugman, Goeric Huybrechts +2

We present an approach to synthesize whisper by applying a handcrafted signal processing recipe and Voice Conversion (VC) techniques to convert normally phonated speech to whispere…