3 papers
eess.AS2026
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
Zhichao Wang, Tao Li, Wenshuo Ge +3
Recent progress of voice conversion~(VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized model…
eess.AS2026
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
Yingying Gao, Shilei Zhang, Runyan Yang +2
Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising…
cs.SD2026
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
Tao Li, Wenshuo Ge, Zhichao Wang +6
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-…