2 papers
cs.CL2023
A Comparative Analysis of Pretrained Language Models for Text-to-Speech
Marcel Granero-Moya, Penny Karanasou, Sri Karlapati +4
State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech. However, while PLMs have b…
eess.AS2023
eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer
Ammar Abbas, Sri Karlapati, Bastian Schnell +7
We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-gr…