collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

Chunyu Qiang, Xiaopeng Wang, Kang Yin +11

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous…

eess.AS2025

SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization

Wenxi Chen, Xinsheng Wang, Ruiqi Yan +9

Speech codecs that convert continuous speech signals into discrete tokens have become essential for speech language models. However, existing codecs struggle to balance high-qualit…

eess.AS2025

AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook

Yushen Chen, Kai Hu, Long Zhou +4

We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, musi…

eess.AS2025

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Qixi Zheng, Yushen Chen, Zhikang Niu +4

Flow-matching-based text-to-speech (TTS) models, such as Voicebox, E2 TTS, and F5-TTS, have attracted significant attention in recent years. These models require multiple sampling…

eess.AS2025

Towards Flow-Matching-based TTS without Classifier-Free Guidance

Yuzhe Liang, Wenzhe Liu, Chunyu Qiang +7

Flow matching has demonstrated strong generative capabilities and has become a core component in modern Text-to-Speech (TTS) systems. To ensure high-quality speech synthesis, Class…

eess.AS2024

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Yushen Chen, Zhikang Niu, Ziyang Ma +5

This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as du…