23 citations · 24 across the 3 of their papers we have counts for
3 papers
Investigating self-supervised features for expressive, multilingual voice conversion
Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer +7
Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between d…
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Mateusz Łajszczak, Guillermo Cámbara, Yang Li +16
We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TT…
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6
Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…