activity
20152023
most citedLeveraging Word Embeddings for Spoken Document Summarization

4 citations · 9 across the 5 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2019

Distributed Microphone Speech Enhancement based on Deep Learning

Syu-Siang Wang, Yu-You Liang, Jeih-weih Hung +3

Speech-related applications deliver inferior performance in complex noise environments. Therefore, this study primarily addresses this problem by introducing speech-enhancement (SE…

eess.AS2019

Generalization of Spectrum Differential based Direct Waveform Modification for Voice Conversion

Wen-Chin Huang, Yi-Chiao Wu, Kazuhiro Kobayashi +6

We present a modification to the spectrum differential based direct waveform modification for voice conversion (DIFFVC) so that it can be directly applied as a waveform generation…

eess.AS20192 cited

Investigation of F0 conditioning and Fully Convolutional Networks in Variational Autoencoder based Voice Conversion

Wen-Chin Huang, Yi-Chiao Wu, Chen-Chou Lo +6

In this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship betw…

eess.AS2018

Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition

Yih-Liang Shen, Chao-Yuan Huang, Syu-Siang Wang +3

Conventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-opt…

eess.AS2018

Refined WaveNet Vocoder for Variational Autoencoder Based Voice Conversion

Wen-Chin Huang, Yi-Chiao Wu, Hsin-Te Hwang +6

This paper presents a refinement framework of WaveNet vocoders for variational autoencoder (VAE) based voice conversion (VC), which reduces the quality distortion caused by the mis…

eess.AS2018

Voice Conversion Based on Cross-Domain Features Using Variational Auto Encoders

Wen-Chin Huang, Hsin-Te Hwang, Yu-Huai Peng +2

An effective approach to non-parallel voice conversion (VC) is to utilize deep neural networks (DNNs), specifically variational auto encoders (VAEs), to model the latent structure…