activity
20242026
collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…

cs.SD2025

Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control

Masato Murata, Koichi Miyazaki, Tomoki Koriyama

Cross-speaker emotion intensity control aims to generate emotional speech of a target speaker with desired emotion intensities using only their neutral speech. A recently proposed…

cs.SD2025

Eigenvoice Synthesis based on Model Editing for Speaker Generation

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine…

cs.SD2024

Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition

Yoshiki Masuyama, Koichi Miyazaki, Masato Murata

Selective state space models (SSMs) represented by Mamba have demonstrated their computational efficiency and promising outcomes in various tasks, including automatic speech recogn…

cs.SD2024

An Attribute Interpolation Method in Speech Synthesis by Model Merging

Masato Murata, Koichi Miyazaki, Tomoki Koriyama

With the development of speech synthesis, recent research has focused on challenging tasks, such as speaker generation and emotion intensity control. Attribute interpolation is a c…

cs.SD2024

Exploring the Capability of Mamba in Speech Applications

Koichi Miyazaki, Yoshiki Masuyama, Masato Murata

This paper explores the capability of Mamba, a recently proposed architecture based on state space models (SSMs), as a competitive alternative to Transformer-based models. In the s…