activity
20182022
most citedVoice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion

34 citations · 49 across the 8 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2020

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

Jennifer Williams, Yi Zhao, Erica Cooper +1

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not…

eess.AS20202 cited

Predictions of Subjective Ratings and Spoofing Assessments of Voice Conversion Challenge 2020 Submissions

Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang +5

The Voice Conversion Challenge 2020 is the third edition under its flagship that promotes intra-lingual semiparallel and cross-lingual voice conversion (VC). While the primary eval…

eess.AS202034 cited

Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion

Yi Zhao, Wen-Chin Huang, Xiaohai Tian +5

The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organiz…

eess.AS20203 cited

Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

Yi Zhao, Haoyu Li, Cheng-I Lai +3

Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without super…

eess.AS20191 cited

Transferring neural speech waveform synthesizers to musical instrument sounds generation

Yi Zhao, Xin Wang, Lauri Juvela +1

Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different meth…

eess.AS20191 cited

Does the Lombard Effect Improve Emotional Communication in Noise? - Analysis of Emotional Speech Acted in Noise -

Yi Zhao, Atsushi Ando, Shinji Takaki +2

Speakers usually adjust their way of talking in noisy environments involuntarily for effective communication. This adaptation is known as the Lombard effect. Although speech accomp…