5 citations · 5 across the 6 of their papers we have counts for
1 paper · 1 filter
Shivam Mehta, Harm Lameris, Rajiv Punmiya +3
Converting input symbols to output audio in TTS requires modelling the durations of speech sounds. Leading non-autoregressive (NAR) TTS models treat duration modelling as a regress…