23 citations · 26 across the 9 of their papers we have counts for
6 papers · 1 filter
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6
Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…
Controllable Emphasis with zero data for text-to-speech
Arnaud Joly, Marco Nicolis, Ekaterina Peterova +11
We present a scalable method to produce high quality emphasis for text-to-speech (TTS) that does not require recordings or annotations. Many TTS models include a phoneme duration m…
Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody
Peter Makarov, Ammar Abbas, Mateusz Łajszczak +5
Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs…
CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer
Sri Karlapati, Penny Karanasou, Mateusz Lajszczak +7
In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually…
Distribution augmentation for low-resource expressive text-to-speech
Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8
This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…
Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech
Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek +2
This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and…