activity
20242026
most citedSpeech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis

2 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2026

Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

Kenichi Fujita, Yusuke Ijima

Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system ge…

cs.SD2025

Multi-interaction TTS toward professional recording reproduction

Hiroki Kanagawa, Kenichi Fujita, Aya Watanabe +1

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is…

cs.SD2024

Lightweight Zero-shot Text-to-Speech with Mixture of Adapters

Kenichi Fujita, Takanori Ashihara, Marc Delcroix +1

The advancements in zero-shot text-to-speech (TTS) methods, based on large-scale models, have demonstrated high fidelity in reproducing speaker characteristics. However, these mode…

cs.SD20242 cited

Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis

Kenichi Fujita, Atsushi Ando, Yusuke Ijima

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essentia…

cs.SD2024

Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters

Kenichi Fujita, Hiroshi Sato, Takanori Ashihara +4

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce sp…