1 paper
Rem Hida, Masaki Hamada, Chie Kamada +3
Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information fr…