most citedVARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention

15 citations · 40 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS20213 cited

DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion

Songxiang Liu, Yuewen Cao, Dan Su +1

Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and exp…

cs.SD202115 cited

VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention

Peng Liu, Yuewen Cao, Songxiang Liu +4

This paper proposes VARA-TTS, a non-autoregressive (non-AR) text-to-speech (TTS) model using a very deep Variational Autoencoder (VDVAE) with Residual Attention mechanism, which re…

eess.AS2020

FastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulation

Songxiang Liu, Yuewen Cao, Na Hu +2

This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than r…

eess.AS20202 cited

Transferring Source Style in Non-Parallel Voice Conversion

Songxiang Liu, Yuewen Cao, Shiyin Kang +5

Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information. Most VC approaches ignore modeling of the sp…

eess.AS202011 cited

Multi-Target Emotional Voice Conversion With Neural Vocoders

Songxiang Liu, Yuewen Cao, Helen Meng

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one em…

eess.AS20209 cited

Emotional Voice Conversion With Cycle-consistent Adversarial Network

Songxiang Liu, Yuewen Cao, Helen Meng

Emotional Voice Conversion, or emotional VC, is a technique of converting speech from one emotion state into another one, keeping the basic linguistic information and speaker ident…