activity
20242026
collaborators

12 papers

cs.SD2026

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4

Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…

cs.SD2026

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

Xiang Li, Yixuan Zhou, Jingran Xie +2

Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Mo…

cs.SD2026

VoxCPM2 Technical Report

Yixuan Zhou, Guoyang Zeng, Xin Liu +15

We present VoxCPM2, a https://info.arxiv.org/help/prep#abstractsfully open-source multilingual and controllable speech generation foundation model that extends the hierarchical dif…

eess.AS2026

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

Zhisheng Zhang, Xiang Li, Yixuan Zhou +3

Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details…

cs.SD2026

A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding

Runchuan Ye, Yixuan Zhou, Renjie Yu +6

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which seve…

cs.SD2025

In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion

Jiawei Jin, Zhihan Yang, Yixuan Zhou +1

We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and envir…