10 papers
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
Yi Gu, Yanqing Liu, Chen Yang +1
Text-to-audio (T2A) generation has advanced considerably in recent years, yet existing methods continue to face challenges in accurately rendering complex text prompts, particularl…
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
Hanchen Pei, Shujie Liu, Yanqing Liu +5
Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody,…
Autoregressive Speech Synthesis without Vector Quantization
Lingwei Meng, Long Zhou, Shujie Liu +9
We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram f…
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Ze Yuan, Yanqing Liu, Shujie Liu +1
Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience an…
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker +10
This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-a…
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
Jiaqi Li, Dongmei Wang, Xiaofei Wang +13
Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the…