9 papers
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong +4
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and ty…
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech
Wei Xue, Junlan Feng, Shilei Zhang +9
Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rel…
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
Zhichao Wang, Tao Li, Wenshuo Ge +3
Recent progress of voice conversion~(VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized model…
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
Yingying Gao, Shilei Zhang, Runyan Yang +2
Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising…
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
Tao Li, Wenshuo Ge, Zhichao Wang +6
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-…
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
Runyan Yang, Yuke Si, Yingying Gao +3
While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well…