3 papers
cs.SD2022
Speech Enhancement with Fullband-Subband Cross-Attention Network
Jun Chen, Wei Rao, Zilin Wang +5
FullSubNet has shown its promising performance on speech enhancement by utilizing both fullband and subband information. However, the relationship between fullband and subband in F…
eess.AS2022
Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using -VAE
Hui Lu, Disong Wang, Xixin Wu +3
We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot c…
cs.SD2022
A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS
Haohan Guo, Fenglong Xie, Frank K. Soong +2
We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is us…