12 citations · 12 across the 3 of their papers we have counts for
3 papers
Diversity-based core-set selection for text-to-speech with linguistic and acoustic features
Kentaro Seki, Shinnosuke Takamichi, Takaaki Saeki +1
This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed f…
Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech
Dong Yang, Tomoki Koriyama, Yuki Saito +3
Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhyth…
JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification
Shinnosuke Takamichi, Ludwig Kürzinger, Takaaki Saeki +2
In this paper, we construct a new Japanese speech corpus called "JTubeSpeech." Although recent end-to-end learning requires large-size speech corpora, open-sourced such corpora for…