activity
20192026
most citedMulti-speaker Text-to-speech Synthesis Using Deep Gaussian Processes

3 citations · 5 across the 11 of their papers we have counts for

collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2026

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…

cs.SD2026

Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement

Junya Koguchi, Tomoki Koriyama

Instantaneous pitch estimation plays an important role in analyzing steep pitch variations such as speech prosody and singing techniques. Conventional approaches estimate instantan…

cs.SD2026

Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection

Junya Koguchi, Tomoki Koriyama

The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a princi…

cs.SD2025

Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control

Masato Murata, Koichi Miyazaki, Tomoki Koriyama

Cross-speaker emotion intensity control aims to generate emotional speech of a target speaker with desired emotion intensities using only their neutral speech. A recently proposed…

cs.SD2025

Eigenvoice Synthesis based on Model Editing for Speaker Generation

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine…

cs.SD2024

An Attribute Interpolation Method in Speech Synthesis by Model Merging

Masato Murata, Koichi Miyazaki, Tomoki Koriyama

With the development of speech synthesis, recent research has focused on challenging tasks, such as speaker generation and emotion intensity control. Attribute interpolation is a c…