activity
20172022
most citedStatistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks

12 citations · 18 across the 8 of their papers we have counts for

collaborators

11 papers

cs.SD2022

Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +2

In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that ai…

cs.SD2022

Multi-Task Adversarial Training Algorithm for Multi-Speaker Neural Text-to-Speech

Yusuke Nakai, Yuki Saito, Kenta Udagawa +1

We propose a novel training algorithm for a multi-speaker neural text-to-speech (TTS) model based on multi-task adversarial training. A conventional generative adversarial network…

cs.HC2021

HumanACGAN: conditional generative adversarial network with human-based auxiliary classifier and its evaluation in phoneme perception

Yota Ueda, Kazuki Fujii, Yuki Saito +3

We propose a conditional generative adversarial network (GAN) incorporating humans' perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a rea…

cs.SD2020

Lifter Training and Sub-band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials

Takaaki Saeki, Yuki Saito, Shinnosuke Takamichi +1

In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials…

cs.SD2019

HumanGAN: generative adversarial network with human-based discriminator and its evaluation in speech perception modeling

Kazuki Fujii, Yuki Saito, Shinnosuke Takamichi +2

We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribut…

cs.SD2019

JVS corpus: free Japanese multi-speaker voice corpus

Shinnosuke Takamichi, Kentaro Mitsui, Yuki Saito +3

Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we ar…