activity
20192024
most citedBASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

23 citations · 26 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2024★ 1 cited

Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…

eess.AS2023

Controllable Emphasis with zero data for text-to-speech

Arnaud Joly, Marco Nicolis, Ekaterina Peterova +11

We present a scalable method to produce high quality emphasis for text-to-speech (TTS) that does not require recordings or annotations. Many TTS models include a phoneme duration m…

eess.AS2022

Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody

Peter Makarov, Ammar Abbas, Mateusz Łajszczak +5

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs…

eess.AS2022

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

Sri Karlapati, Penny Karanasou, Mateusz Lajszczak +7

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually…

eess.AS2022

Distribution augmentation for low-resource expressive text-to-speech

Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…

eess.AS2019

Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech

Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek +2

This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and…