most citedAlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment

4 citations · 16 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS20214 cited

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

Zhen Zeng, Jianzong Wang, Ning Cheng +1

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use o…

eess.AS20202 cited

GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis

Aolan Sun, Jianzong Wang, Ning Cheng +4

This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relat…

cs.SD20202 cited

MelGlow: Efficient Waveform Generative Network Based on Location-Variable Convolution

Zhen Zeng, Jianzong Wang, Ning Cheng +1

Recent neural vocoders usually use a WaveNet-like network to capture the long-term dependencies of the waveform, but a large number of parameters are required to obtain good modeli…

eess.AS20201 cited

Prosody Learning Mechanism for Speech Synthesis System Without Text Length Limit

Zhen Zeng, Jianzong Wang, Ning Cheng +1

Recent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of pr…

eess.AS20204 cited

AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment

Zhen Zeng, Jianzong Wang, Ning Cheng +2

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-…

eess.AS20203 cited

GraphTTS: graph-to-sequence modelling in neural text-to-speech

Aolan Sun, Jianzong Wang, Ning Cheng +3

This paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs co…