activity
20182022
most citedVoice Activity Detection: Merging Source and Filter-based Information

82 citations · 231 across the 39 of their papers we have counts for

collaborators

53 papers

eess.AS2022

Distribution augmentation for low-resource expressive text-to-speech

Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…

eess.AS2021

Multi-Scale Spectrogram Modelling for Neural Text-to-Speech

Ammar Abbas, Bajibabu Bollepalli, Alexis Moinet +6

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrog…

cs.SD2021

Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant Environments

Alejandro Mottini, Jaime Lorenzo-Trueba, Sri Vishnu Kumar Karlapati +1

Voice Conversion (VC) is a technique that aims to transform the non-linguistic information of a source utterance to change the perceived identity of the speaker. While there is a r…

eess.AS2021

A learned conditional prior for the VAE acoustic space of a TTS system

Penny Karanasou, Sri Karlapati, Alexis Moinet +5

Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…

eess.AS2021

Weakly-supervised word-level pronunciation error detection in non-native English speech

Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman +2

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not req…

eess.AS20211 cited

Mispronunciation Detection in Non-native (L2) English with Uncertainty Modeling

Daniel Korzekwa, Jaime Lorenzo-Trueba, Szymon Zaporowski +3

A common approach to the automatic detection of mispronunciation in language learning is to recognize the phonemes produced by a student and compare it to the expected pronunciatio…