activity
20172022
most citedPhonetic Posteriorgrams based Many-to-Many Singing Voice Conversion via Adversarial Training

10 citations · 23 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2022

VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion

Disong Wang, Shan Yang, Dan Su +3

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speec…

eess.AS2021

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

Songxiang Liu, Shan Yang, Dan Su +1

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSS…

eess.AS20213 cited

Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis

Jian Cong, Shan Yang, Lei Xie +1

Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectr…

eess.AS2021

Controllable Context-aware Conversational Speech Synthesis

Jian Cong, Shan Yang, Na Hu +3

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interloc…

eess.AS20204 cited

Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial Training

Jian Cong, Shan Yang, Lei Xie +2

Data efficient voice cloning aims at synthesizing target speaker's voice with only a few enrollment samples at hand. To this end, speaker adaptation and speaker encoding are two ty…