most citedEfficient Neural Music Generation

12 citations · 19 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2023

U-Style: Cascading U-nets with Multi-level Speaker and Style Modeling for Zero-Shot Voice Cloning

Tao Li, Zhichao Wang, Xinfa Zhu +4

Zero-shot speaker cloning aims to synthesize speech for any target speaker unseen during TTS system building, given only a single speech reference of the speaker at hand. Although…

cs.SD20231 cited

AudioSR: Versatile Audio Super-resolution at Scale

Haohe Liu, Ke Chen, Qiao Tian +2

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods hav…

eess.AS2023

MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling

Zhichao Wang, Xinsheng Wang, Qicong Xie +4

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice co…

cs.SD20231 cited

DiCLET-TTS: Diffusion Model based Cross-lingual Emotion Transfer for Text-to-Speech -- A Study between English and Mandarin

Tao Li, Chenxu Hu, Jian Cong +5

While the performance of cross-lingual TTS based on monolingual corpora has been significantly improved recently, generating cross-lingual speech still suffers from the foreign acc…

cs.SD202312 cited

Efficient Neural Music Generation

Max W. Y. Lam, Qiao Tian, Tang Li +10

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acous…

cs.SD20225 cited

Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech

Zhengxi Liu, Qiao Tian, Chenxu Hu +5

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms direct…