2 papers
eess.AS2024
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
Shivam Mehta, Harm Lameris, Rajiv Punmiya +3
Converting input symbols to output audio in TTS requires modelling the durations of speech sounds. Leading non-autoregressive (NAR) TTS models treat duration modelling as a regress…
eess.AS2023
Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis
Erik Ekstedt, Siyang Wang, Éva Székely +2
Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently prop…