13 citations · 24 across the 5 of their papers we have counts for
13 papers · 1 filter
Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer
Jakub Swiatkowski, Duo Wang, Mikolaj Babianski +5
Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking…
Cross-lingual Prosody Transfer for Expressive Machine Dubbing
Jakub Swiatkowski, Duo Wang, Mikolaj Babianski +3
Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this…
crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
Kazuhiro Kobayashi, Wen-Chin Huang, Yi-Chiao Wu +3
In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based…
The NU Voice Conversion System for the Voice Conversion Challenge 2020: On the Effectiveness of Sequence-to-sequence Models and Autoregressive Neural Vocoders
Wen-Chin Huang, Patrick Lumban Tobing, Yi-Chiao Wu +2
In this paper, we present the voice conversion (VC) systems developed at Nagoya University (NU) for the Voice Conversion Challenge 2020 (VCC2020). We aim to determine the effective…
Quasi-Periodic WaveNet: An Autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing +2
In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the limited pitch controllability of vanilla WaveNet (WN) usin…
A Cyclical Post-filtering Approach to Mismatch Refinement of Neural Vocoder for Text-to-speech Systems
Yi-Chiao Wu, Patrick Lumban Tobing, Kazuki Yasuhara +3
Recently, the effectiveness of text-to-speech (TTS) systems combined with neural vocoders to generate high-fidelity speech has been shown. However, collecting the required training…