activity
20172023
most citedParallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks

178 citations · 270 across the 17 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

cs.SD2020

CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka +1

Non-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus. Recently, cycle-consistent adversarial n…

cs.SD2020

VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics

Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka +2

In this paper, we propose a non-parallel any-to-many voice conversion (VC) method termed VoiceGrad. Inspired by WaveGrad, a recently introduced novel waveform generation method, Vo…

eess.AS2020

X-DC: Explainable Deep Clustering based on Learnable Spectrogram Templates

Chihiro Watanabe, Hirokazu Kameoka

Deep neural networks (DNNs) have achieved substantial predictive performance in various speech processing tasks. Particularly, it has been shown that a monaural speech separation t…

eess.AS2020★ 3 cited

Pretraining Techniques for Sequence-to-Sequence Voice Conversion

Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu +2

Sequence-to-sequence (seq2seq) voice conversion (VC) models are attractive owing to their ability to convert prosody. Nonetheless, without sufficient data, seq2seq VC models can su…

eess.AS2020

Nonparallel Voice Conversion with Augmented Classifier Star Generative Adversarial Networks

Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka +1

We previously proposed a method that allows for nonparallel voice conversion (VC) by using a variant of generative adversarial networks (GANs) called StarGAN. The main features of…

eess.AS2020

Many-to-Many Voice Transformer Network

Hirokazu Kameoka, Wen-Chin Huang, Kou Tanaka +3

This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pit…