28 citations · 29 across the 5 of their papers we have counts for
8 papers
SpeechLMScore: Evaluating speech generation using speech language model
Soumi Maiti, Yifan Peng, Takaaki Saeki +1
While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality…
Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection
Kentaro Seki, Shinnosuke Takamichi, Takaaki Saeki +1
This paper proposes a method for selecting training data for text-to-speech (TTS) synthesis from dark data. TTS models are typically trained on high-quality speech corpora that cos…
Personalized Filled-pause Generation with Group-wise Prediction Models
Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi +1
In this paper, we propose a method to generate personalized filled pauses (FPs) with group-wise prediction models. Compared with fluent text generation, disfluent text generation h…
vTTS: visual-text to speech
Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi +2
This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into…
ESPnet2-TTS: Extending the Edge of TTS Research
Tomoki Hayashi, Ryuichi Yamamoto, Takenori Yoshimura +7
This paper describes ESPnet2-TTS, an end-to-end text-to-speech (E2E-TTS) toolkit. ESPnet2-TTS extends our earlier version, ESPnet-TTS, by adding many new features, including: on-th…
Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
Incremental text-to-speech (TTS) synthesis generates utterances in small linguistic units for the sake of real-time and low-latency applications. We previously proposed an incremen…