activity
20222024
most citedBASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

23 citations · 24 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG202423 cited

BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Mateusz Łajszczak, Guillermo Cámbara, Yang Li +16

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TT…

cs.LG20241 cited

Energy-conserving equivariant GNN for elasticity of lattice architected metamaterials

Ivan Grega, Ilyes Batatia, Gábor Csányi +2

Lattices are architected metamaterials whose properties strongly depend on their geometrical design. The analogy between lattices and graphs enables the use of graph neural network…

cs.CL2023

A Comparative Analysis of Pretrained Language Models for Text-to-Speech

Marcel Granero-Moya, Penny Karanasou, Sri Karlapati +4

State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech. However, while PLMs have b…

eess.AS2022

Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody

Peter Makarov, Ammar Abbas, Mateusz Łajszczak +5

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs…

eess.AS2022

Expressive, Variable, and Controllable Duration Modelling in TTS

Ammar Abbas, Thomas Merritt, Alexis Moinet +5

Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to rely…

eess.AS2022

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

Sri Karlapati, Penny Karanasou, Mateusz Lajszczak +7

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually…