activity
20192026
most citedMulti-speaker Text-to-speech Synthesis Using Deep Gaussian Processes

3 citations · 5 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2025

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model

Dong Yang, Yuki Saito, Takaaki Saeki +4

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker emb…

eess.AS2025

Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Tomoki Koriyama

This paper proposes a model for automatic prosodic label annotation, where the predicted labels can be used for training a prosody-controllable text-to-speech model. The proposed m…

eess.AS2024

VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features

Tomoki Koriyama

This paper presents an accurate phoneme alignment model that aims for speech analysis and video content creation. We propose a variational autoencoder (VAE)-based alignment model i…

eess.AS20203 cited

Multi-speaker Text-to-speech Synthesis Using Deep Gaussian Processes

Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari

Multi-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model. Although many approaches using deep neural networks (DNNs) have been propo…

eess.AS2020

Utterance-level Sequential Modeling For Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit

Tomoki Koriyama, Hiroshi Saruwatari

This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively wit…