papers

Publications (10)

cs.SD2023

Accent-VITS:accent transfer for end-to-end TTS

Linhan Ma, Yongmao Zhang, Xinfa Zhu +4

Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker tim…

eess.AS2025

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech

Linhan Ma, Dake Guo, He Wang +2

Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in…

cs.SD2026

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24

Qwen-Music is a large language model‑based system that generates high‑fidelity songs with vocals from text prompts or re‑imagines existing tracks, using a semantic token representa…

#music generation#text-to-music#cover song generation#semantic tokenization
eess.AS2026

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

Huakang Chen, Jingbin Hu, Liumeng Xue +12

Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due t…

eess.AS2024

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Linhan Ma, Dake Guo, Kun Song +7

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we pr…

cs.SD2024

Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Linhan Ma, Xinfa Zhu, Yuanjun Lv +5

Zero-shot voice conversion (VC) aims to transform source speech into arbitrary unseen target voice while keeping the linguistic content unchanged. Recent VC methods have made signi…