most citedVALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

8 citations · 11 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Autoregressive Speech Synthesis without Vector Quantization

Lingwei Meng, Long Zhou, Shujie Liu +9

We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram f…

cs.CL20248 cited

VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Sanyuan Chen, Shujie Liu, Long Zhou +6

This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity fo…

cs.CL20241 cited

VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Bing Han, Long Zhou, Shujie Liu +7

With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis…

cs.CL2024

WavLLM: Towards Robust and Adaptive Speech Large Language Model

Shujie Hu, Long Zhou, Shujie Liu +9

The recent advancements in large language models (LLMs) have revolutionized the field of natural language processing, progressively broadening their scope to multimodal perception…

cs.CL20232 cited

Boosting Large Language Model for Speech Synthesis: An Empirical Study

Hongkun Hao, Long Zhou, Shujie Liu +4

Large language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as spee…