activity
20182020
collaborators

6 papers

eess.AS2020

Unsupervised Representation Disentanglement using Cross Domain Features and Adversarial Learning in Variational Autoencoder based Voice Conversion

Wen-Chin Huang, Hao Luo, Hsin-Te Hwang +4

An effective approach for voice conversion (VC) is to disentangle linguistic content from other components in the speech signal. The effectiveness of variational autoencoder (VAE)…

eess.AS2019

ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech

Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37

Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…

eess.AS2019

Generalization of Spectrum Differential based Direct Waveform Modification for Voice Conversion

Wen-Chin Huang, Yi-Chiao Wu, Kazuhiro Kobayashi +6

We present a modification to the spectrum differential based direct waveform modification for voice conversion (DIFFVC) so that it can be directly applied as a waveform generation…

eess.AS2018

Refined WaveNet Vocoder for Variational Autoencoder Based Voice Conversion

Wen-Chin Huang, Yi-Chiao Wu, Hsin-Te Hwang +6

This paper presents a refinement framework of WaveNet vocoders for variational autoencoder (VAE) based voice conversion (VC), which reduces the quality distortion caused by the mis…

eess.AS2018

Voice Conversion Based on Cross-Domain Features Using Variational Auto Encoders

Wen-Chin Huang, Hsin-Te Hwang, Yu-Huai Peng +2

An effective approach to non-parallel voice conversion (VC) is to utilize deep neural networks (DNNs), specifically variational auto encoders (VAEs), to model the latent structure…

cs.SD2018

Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang +1

Nowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech…