2 papers
eess.AS2026
Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis
Biel Tura Vecino, Yoach Lacombe, Julian Weber +4
Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and un…
eess.AS2024
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Edresson Casanova, Kelly Davis, Eren Gölge +8
Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they…