activity
20182022
most citedUniversal Neural Vocoding with Parallel WaveNet

2 citations · 7 across the 9 of their papers we have counts for

collaborators

11 papers

eess.AS2022

Remap, warp and attend: Non-parallel many-to-many accent conversion with Normalizing Flows

Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa +6

Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intona…

eess.AS20221 cited

Automated detection of pronunciation errors in non-native English speech employing deep learning

Daniel Korzekwa

Despite significant advances in recent years, the existing Computer-Assisted Pronunciation Training (CAPT) methods detect pronunciation errors with a relatively low accuracy (preci…

eess.AS20222 cited

Text-free non-parallel many-to-many voice conversion using normalising flows

Thomas Merritt, Abdelhamid Ezzerg, Piotr Biliński +4

Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all…

eess.AS2021

Enhancing audio quality for expressive Neural Text-to-Speech

Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz +7

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recor…

cs.SD20211 cited

Non-Autoregressive TTS with Explicit Duration Modelling for Low-Resource Highly Expressive Speech

Raahil Shah, Kamil Pokora, Abdelhamid Ezzerg +5

Whilst recent neural text-to-speech (TTS) approaches produce high-quality speech, they typically require a large amount of recordings from the target speaker. In previous work, a 3…

eess.AS2021

Weakly-supervised word-level pronunciation error detection in non-native English speech

Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman +2

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not req…