6 papers
Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang +3
Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion o…
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
Jiacheng Shi, Hongfei Du, Xinyuan Song +3
Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acousti…
MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
Luz Martinez-Lucas, Pravin Mote, Abinay Reddy Naini +2
Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions convey…
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
Chan-Jan Hsu, Liang-Hsuan Tseng, Yi-Cheng Lin +5
Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving attributes like speaker and emotion, se…
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
Jing-Tong Tzeng, Carlos Busso, Chi-Chun Lee
Speech emotion recognition (SER) plays a critical role in building emotion-aware speech systems, but its performance degrades significantly under noisy conditions. Although speech…
Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela +2
Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach th…