3 papers
eess.AS2020
WG-WaveNet: Real-Time High-Fidelity Speech Synthesis without GPU
Po-chun Hsu, Hung-yi Lee
In this paper, we propose WG-WaveNet, a fast, lightweight, and high-quality waveform generation model. WG-WaveNet is composed of a compact flow-based model and a post-filter. The t…
cs.SD2019
Towards Robust Neural Vocoding for Speech Generation: A Survey
Po-chun Hsu, Chun-hsuan Wang, Andy T. Liu +1
Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion. However, when encountering data distribution mismatch betw…
eess.AS2019
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
Andy T. Liu, Shu-wen Yang, Po-Han Chi +2
We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech. Previous spe…