activity
20242026
collaborators

5 papers

cs.SD2026

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…

cs.SD2026

Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens

Kazuki Yamauchi, Masato Murata, Shogo Seki

Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech…

cs.SD2025

Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control

Masato Murata, Koichi Miyazaki, Tomoki Koriyama

Cross-speaker emotion intensity control aims to generate emotional speech of a target speaker with desired emotion intensities using only their neutral speech. A recently proposed…

cs.SD2025

Eigenvoice Synthesis based on Model Editing for Speaker Generation

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine…

cs.SD2024

Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition

Yoshiki Masuyama, Koichi Miyazaki, Masato Murata

Selective state space models (SSMs) represented by Mamba have demonstrated their computational efficiency and promising outcomes in various tasks, including automatic speech recogn…