activity
20182022
most citedFeature reinforcement with word embedding and parsing information in neural TTS

13 citations · 46 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20214 cited

Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation

Fengpeng Yue, Yan Deng, Lei He +1

Machine Speech Chain, which integrates both end-to-end (E2E) automatic speech recognition (ASR) and text-to-speech (TTS) into one circle for joint training, has been proven to be e…

cs.CL2021

Multilingual Byte2Speech Models for Scalable Low-resource Speech Synthesis

Mutian He, Jingzhou Yang, Lei He +1

To scale neural speech synthesis to various real-world languages, we present a multilingual end-to-end framework that maps byte inputs to spectrograms, thus allowing arbitrary inpu…

cs.CL20196 cited

A New GAN-based End-to-End TTS Training Algorithm

Haohan Guo, Frank K. Soong, Lei He +1

End-to-end, autoregressive model-based TTS has shown significant performance improvements over the conventional one. However, the autoregressive module training is affected by the…

cs.CL20192 cited

Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS

Haohan Guo, Frank K. Soong, Lei He +1

The end-to-end TTS, which can predict speech directly from a given sequence of graphemes or phonemes, has shown improved performance over the conventional TTS. However, its predict…

cs.CL20183 cited

Recurrent Neural Networks with Pre-trained Language Model Embedding for Slot Filling Task

Liang Qiu, Yuanyi Ding, Lei He

In recent years, Recurrent Neural Networks (RNNs) based models have been applied to the Slot Filling problem of Spoken Language Understanding and achieved the state-of-the-art perf…

cs.CL2018

Learning latent representations for style control and transfer in end-to-end speech synthesis

Ya-Jie Zhang, Shifeng Pan, Lei He +1

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised mann…