activity
20232026
most citedVALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

10 citations · 17 across the 25 of their papers we have counts for

collaborators

31 papers

cs.SD2026

Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

Qingyu Liu, Rixi Xu, Yushen Chen +11

Zero-shot text-to-speech (TTS) can clone a speaker's voice from a short audio prompt, yet most TTS systems still require the audio prompt transcript during inference. This dependen…

cs.SD2026

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Ziyang Ma, Zhikang Niu, Wenming Tu +30

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To sup…

eess.AS2026

GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model

Guanrou Yang, Tian Tan, Qian Chen +8

Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE…

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +35

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

cs.SD2026

MMAE: A Massive Multitask Audio Editing Benchmark

Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…

eess.AS2026

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Wenxi Chen, Dongya Jia, Yushen Chen +11

Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve gener…