12 papers
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech
Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4
Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment
Xiang Li, Yixuan Zhou, Jingran Xie +2
Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Mo…
VoxCPM2 Technical Report
Yixuan Zhou, Guoyang Zeng, Xin Liu +15
We present VoxCPM2, a https://info.arxiv.org/help/prep#abstractsfully open-source multilingual and controllable speech generation foundation model that extends the hierarchical dif…
LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
Zhisheng Zhang, Xiang Li, Yixuan Zhou +3
Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details…
A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
Runchuan Ye, Yixuan Zhou, Renjie Yu +6
Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which seve…
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
Jiawei Jin, Zhihan Yang, Yixuan Zhou +1
We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and envir…