activity
20172022
most citedParallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks

178 citations · 263 across the 12 of their papers we have counts for

collaborators

28 papers

eess.AS20221 cited

DisC-VC: Disentangled and F0-Controllable Neural Voice Conversion

Chihiro Watanabe, Hirokazu Kameoka

Voice conversion is a task to convert a non-linguistic feature of a given utterance. Since naturalness of speech strongly depends on its pitch pattern, in some applications, it wou…

cs.SD2022

iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform

Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka +1

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vo…

cs.SD20218 cited

FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion

Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko

This paper proposes a non-autoregressive extension of our previously proposed sequence-to-sequence (S2S) model-based voice conversion (VC) methods. S2S model-based VC methods have…

cs.SD20216 cited

StarGAN-based Emotional Voice Conversion for Japanese Phrases

Asuka Moritani, Ryo Ozaki, Shoki Sakamoto +2

This paper shows that StarGAN-VC, a spectral envelope transformation method for non-parallel many-to-many voice conversion (VC), is capable of emotional VC (EVC). Although StarGAN-…

cs.SD20211 cited

MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka +1

Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-…

cs.SD2020

CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka +1

Non-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus. Recently, cycle-consistent adversarial n…