10 citations · 17 across the 25 of their papers we have counts for
31 papers
Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning
Qingyu Liu, Rixi Xu, Yushen Chen +11
Zero-shot text-to-speech (TTS) can clone a speaker's voice from a short audio prompt, yet most TTS systems still require the audio prompt transcript during inference. This dependen…
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Ziyang Ma, Zhikang Niu, Wenming Tu +30
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To sup…
GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Guanrou Yang, Tian Tan, Qian Chen +8
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE…
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +35
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
Wenxi Chen, Dongya Jia, Yushen Chen +11
Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve gener…