23 citations · 24 across the 6 of their papers we have counts for
6 papers
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Mateusz Łajszczak, Guillermo Cámbara, Yang Li +16
We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TT…
Energy-conserving equivariant GNN for elasticity of lattice architected metamaterials
Ivan Grega, Ilyes Batatia, Gábor Csányi +2
Lattices are architected metamaterials whose properties strongly depend on their geometrical design. The analogy between lattices and graphs enables the use of graph neural network…
A Comparative Analysis of Pretrained Language Models for Text-to-Speech
Marcel Granero-Moya, Penny Karanasou, Sri Karlapati +4
State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech. However, while PLMs have b…
Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody
Peter Makarov, Ammar Abbas, Mateusz Łajszczak +5
Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs…
Expressive, Variable, and Controllable Duration Modelling in TTS
Ammar Abbas, Thomas Merritt, Alexis Moinet +5
Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to rely…
CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer
Sri Karlapati, Penny Karanasou, Mateusz Lajszczak +7
In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually…