collaborators

9 papers

cs.SD2025

LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models

Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka +1

Previously, we introduced VoiceGrad, a nonparallel voice conversion (VC) technique enabling mel-spectrogram conversion from source to target speakers using a score-based diffusion…

cs.SD2025

Vocoder-Projected Feature Discriminator

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka +1

In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and…

cs.SD2025

FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka +1

A diffusion-based voice conversion (VC) model (e.g., VoiceGrad) can achieve high speech quality and speaker similarity; however, its conversion process is slow owing to iterative s…

cs.SD2025

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles

Yuto Kondo, Hirokazu Kameoka, Kou Tanaka +1

We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate…

cs.SD2025

Learning to assess subjective impressions from speech

Yuto Kondo, Hirokazu Kameoka, Kou Tanaka +2

We tackle a new task of training neural network models that can assess subjective impressions conveyed through speech and assign scores accordingly, inspired by the work on automat…

cs.SD2025

Selecting N-lowest scores for training MOS prediction models

Yuto Kondo, Hirokazu Kameoka, Kou Tanaka +1

The automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models ha…