activity
20172022
most citedJSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

88 citations · 140 across the 19 of their papers we have counts for

collaborators

27 papers

cs.SD2022

Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts

Detai Xin, Sharath Adavanne, Federico Ang +3

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context…

cs.SD2022

Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection

Kentaro Seki, Shinnosuke Takamichi, Takaaki Saeki +1

This paper proposes a method for selecting training data for text-to-speech (TTS) synthesis from dark data. TTS models are typically trained on high-quality speech corpora that cos…

cs.SD2022

Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +2

In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that ai…

cs.SD2022

Visual onoma-to-wave: environmental sound synthesis from visual onomatopoeias and sound-source images

Hien Ohnaka, Shinnosuke Takamichi, Keisuke Imoto +3

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e.…

cs.SD2022

Speaking-Rate-Controllable HiFi-GAN Using Feature Interpolation

Detai Xin, Shinnosuke Takamichi, Takuma Okamoto +2

This paper presents a speaking-rate-controllable HiFi-GAN neural vocoder. Original HiFi-GAN is a high-fidelity, computationally efficient, and tiny-footprint neural vocoder. We att…

cs.SD2022

Personalized Filled-pause Generation with Group-wise Prediction Models

Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi +1

In this paper, we propose a method to generate personalized filled pauses (FPs) with group-wise prediction models. Compared with fluent text generation, disfluent text generation h…