activity
20222024
most citedBASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

23 citations · 24 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG202423 cited

BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Mateusz Łajszczak, Guillermo Cámbara, Yang Li +16

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TT…

eess.AS20241 cited

Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…

eess.AS2023

Controllable Emphasis with zero data for text-to-speech

Arnaud Joly, Marco Nicolis, Ekaterina Peterova +11

We present a scalable method to produce high quality emphasis for text-to-speech (TTS) that does not require recordings or annotations. Many TTS models include a phoneme duration m…

eess.AS2022

Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody

Peter Makarov, Ammar Abbas, Mateusz Łajszczak +5

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs…

eess.AS2022

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

Sri Karlapati, Penny Karanasou, Mateusz Lajszczak +7

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually…