activity
20182022
most citedTowards High-Quality Neural TTS for Low-Resource Languages by Learning Compact Speech Representations

2 citations · 3 across the 4 of their papers we have counts for

collaborators

5 papers

cs.SD20222 cited

Towards High-Quality Neural TTS for Low-Resource Languages by Learning Compact Speech Representations

Haohan Guo, Fenglong Xie, Xixin Wu +2

This paper aims to enhance low-resource TTS by reducing training data requirements using compact speech representations. A Multi-Stage Multi-Codebook (MSMC) VQ-GAN is trained to le…

cs.SD2022

A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

Haohan Guo, Fenglong Xie, Frank K. Soong +2

We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is us…

cs.CL2021

Nana-HDR: A Non-attentive Non-autoregressive Hybrid Model for TTS

Shilun Lin, Wenchao Su, Li Meng +3

This paper presents Nana-HDR, a new non-attentive non-autoregressive model with hybrid Transformer-based Dense-fuse encoder and RNN-based decoder for TTS. It mainly consists of thr…

cs.CL20211 cited

Triple M: A Practical Text-to-speech Synthesis System With Multi-guidance Attention And Multi-band Multi-time LPCNet

Shilun Lin, Fenglong Xie, Li Meng +2

In this work, a robust and efficient text-to-speech (TTS) synthesis system named Triple M is proposed for large-scale online application. The key components of Triple M are: 1) A s…

eess.AS2018

LP-WaveNet: Linear Prediction-based WaveNet Speech Synthesis

Min-Jae Hwang, Frank Soong, Eunwoo Song +3

We propose a linear prediction (LP)-based waveform generation method via WaveNet vocoding framework. A WaveNet-based neural vocoder has significantly improved the quality of parame…