6 papers
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
Yingying Gao, Shilei Zhang, Runyan Yang +2
Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising…
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
Tao Li, Wenshuo Ge, Zhichao Wang +6
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-…
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
Runyan Yang, Yuke Si, Yingying Gao +3
While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well…
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
Yuke Si, Runyan Yang, Yingying Gao +3
Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared archi…
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
Xiaoxue Luo, Jinwei Huang, Runyan Yang +4
Universal audio codecs learn entangled representations across audio types, whereas some specific codecs offer decoupled representations but are limited to speech. Real-world audio,…
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
Jiaxuan Liu, Zhaoci Liu, Yajun Hu +3
Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffSty…