Showing eess.ASShow all
2 papers · 1 filter
eess.AS2023
Unified speech and gesture synthesis using flow matching
Shivam Mehta, Ruibo Tu, Simon Alexanderson +3
As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviou…
eess.AS2023
Matcha-TTS: A fast TTS architecture with conditional flow matching
Shivam Mehta, Ruibo Tu, Jonas Beskow +2
We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an OD…