most citedMake-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

47 citations · 121 across the 11 of their papers we have counts for

collaborators

15 papers

eess.AS2024

Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…

cs.SD2024

Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Haohan Guo, Fenglong Xie, Dongchao Yang +2

The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by ``recency bias", CLM lacks sufficient attentio…

cs.SD2024

SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis

Haohan Guo, Fenglong Xie, Kun Xie +4

The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered…

cs.SD2024

SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Dongchao Yang, Rongjie Huang, Yuanyuan Wang +5

Scaling Text-to-speech (TTS) to large-scale datasets has been demonstrated as an effective method for improving the diversity and naturalness of synthesized speech. At the high lev…

eess.AS202420 cited

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Zeqian Ju, Yuancheng Wang, Kai Shen +16

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intric…

eess.AS2023

PromptTTS 2: Describing and Generating Voices with Text Prompt

Yichong Leng, Zhifang Guo, Kai Shen +12

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods rel…