activity
20192022
most citedMulti-speaker Text-to-speech Synthesis Using Deep Gaussian Processes

3 citations · 5 across the 4 of their papers we have counts for

collaborators

5 papers

cs.SD20221 cited

Structured State Space Decoder for Speech Recognition and Synthesis

Koichi Miyazaki, Masato Murata, Tomoki Koriyama

Automatic speech recognition (ASR) systems developed in recent years have shown promising results with self-attention models (e.g., Transformer and Conformer), which are replacing…

eess.AS20203 cited

Multi-speaker Text-to-speech Synthesis Using Deep Gaussian Processes

Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari

Multi-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model. Although many approaches using deep neural networks (DNNs) have been propo…

eess.AS2020

Utterance-level Sequential Modeling For Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit

Tomoki Koriyama, Hiroshi Saruwatari

This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively wit…

cs.SD2019

JVS corpus: free Japanese multi-speaker voice corpus

Shinnosuke Takamichi, Kentaro Mitsui, Yuki Saito +3

Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we ar…

cs.SD20191 cited

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

Hiroki Tamaru, Yuki Saito, Shinnosuke Takamichi +2

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice sy…