1 paper · 1 filter
Mattias Cross, Minghui Zhao, Anton Ragni
Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this leng…