activity
20222024
most citedComputer-assisted Pronunciation Training -- Speech synthesis is almost all you need

39 citations · 41 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS20241 cited

Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…

cs.CL2023

Multilingual context-based pronunciation learning for Text-to-Speech

Giulia Comini, Manuel Sam Ribeiro, Fan Yang +2

Phonetic information and linguistic knowledge are an essential component of a Text-to-speech (TTS) front-end. Given a language, a lexicon can be collected offline and Grapheme-to-P…

eess.AS2023

Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech

Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro +10

Neural text-to-speech systems are often optimized on L1/L2 losses, which make strong assumptions about the distributions of the target data space. Aiming to improve those assumptio…

eess.AS2023

Improving grapheme-to-phoneme conversion by learning pronunciations from speech recordings

Manuel Sam Ribeiro, Giulia Comini, Jaime Lorenzo-Trueba

The Grapheme-to-Phoneme (G2P) task aims to convert orthographic input into a discrete phonetic representation. G2P conversion is beneficial to various speech processing application…

eess.AS20221 cited

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation

Giulia Comini, Goeric Huybrechts, Manuel Sam Ribeiro +2

The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming. To overcome these issues, we demonstrate how to bui…

eess.AS202239 cited

Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need

Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman +1

The research community has long studied computer-assisted pronunciation training (CAPT) methods in non-native speech. Researchers focused on studying various model architectures, s…