5 papers
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
Hongfei Du, Jiacheng Shi, Sidi Lu +2
Integrating large language models (LLMs) into text-to-speech (TTS) systems has improved speech expressiveness, yet interpretable emotional control remains challenging. Existing app…
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
Jiacheng Shi, Hongfei Du, Xinyuan Song +3
Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acousti…
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
Jiacheng Shi, Hongfei Du, Yangfan He +2
Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utter…
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
Jiacheng Shi, Hongfei Du, Y. Alicia Hong +1
Speech emotion recognition (SER) with audio-language models (ALMs) remains vulnerable to distribution shifts at test time, leading to performance degradation in out-of-domain scena…
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
Jiacheng Shi, Hongfei Du, Y. Alicia Hong +1
Large audio-language models (LALMs) exhibit strong zero-shot performance across speech tasks but struggle with speech emotion recognition (SER) due to weak paralinguistic modeling…