activity
20192022
most citedMultilingual Neural Machine Translation with Knowledge Distillation

128 citations · 246 across the 20 of their papers we have counts for

collaborators
Showing eess.ASShow all

12 papers · 1 filter

eess.AS202228 cited

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Rongjie Huang, Max W. Y. Lam, Jun Wang +4

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…

eess.AS2022

Learning the Beauty in Songs: Neural Singing Voice Beautifier

Jinglin Liu, Chengxi Li, Yi Ren +2

We are interested in a novel task, singing voice beautifying (SVB). Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, wh…

eess.AS2022

Revisiting Over-Smoothness in Text to Speech

Yi Ren, Xu Tan, Tao Qin +2

Non-autoregressive text to speech (NAR-TTS) models have attracted much attention from both academia and industry due to their fast generation speed. One limitation of NAR-TTS model…

eess.AS2022

ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech

Yi Ren, Ming Lei, Zhiying Huang +4

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted p…

eess.AS20227 cited

MR-SVS: Singing Voice Synthesis with Multi-Reference Encoder

Shoutong Wang, Jinglin Liu, Yi Ren +3

Multi-speaker singing voice synthesis is to generate the singing voice sung by different speakers. To generalize to new speakers, previous zero-shot singing adaptation methods obta…

eess.AS20206 cited

DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling

Chen Zhang, Yi Ren, Xu Tan +5

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many sce…