papers

Publications (42)

eess.AS2025

Can Emotion Fool Anti-spoofing?

Aurosweta Mahapatra, Ismail Rasim Ulgen, Abinay Reddy Naini +2

Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness…

cs.SD2024

Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition

Shreya G. Upadhyay, Ali N. Salman, Carlos Busso +1

Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on ad…

cs.CL2018

End-to-end Audiovisual Speech Activity Detection with Bimodal Recurrent Neural Models

Fei Tao, Carlos Busso

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environ…

cs.MM2019

Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps

Shih-Fu Chang, Alex Hauptmann, Louis-Philippe Morency +18

With the transformative technologies and the rapidly changing global R&D landscape, the multimedia and multimodal community is now faced with many new opportunities and uncertainti…

eess.AS2026

Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition

Jing-Tong Tzeng, Carlos Busso, Chi-Chun Lee

Speech emotion recognition (SER) plays a critical role in building emotion-aware speech systems, but its performance degrades significantly under noisy conditions. Although speech…

cs.CV2022

Driving Anomaly Detection Using Conditional Generative Adversarial Network

Yuning Qiu, Teruhisa Misu, Carlos Busso

Anomaly driving detection is an important problem in advanced driver assistance systems (ADAS). It is important to identify potential hazard scenarios as early as possible to avoid…

cs.CL2026

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

Chan-Jan Hsu, Liang-Hsuan Tseng, Yi-Cheng Lin +5

Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving attributes like speaker and emotion, se…

cs.HC2018

Speech-Driven Expressive Talking Lips with Conditional Sequential Generative Adversarial Networks

Najmeh Sadoughi, Carlos Busso

Articulation, emotion, and personality play strong roles in the orofacial movements. To improve the naturalness and expressiveness of virtual agents (VAs), it is important that we…

cs.SD2026

Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model

Chung-Soo Ahn, Rajib Rana, Sunil Sivadas +2

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more compl…

cs.CL2025

Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition

Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee

Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effect…

eess.AS2025

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Rui Liu, Pu Gao, Jiatian Xi +3

Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consiste…

eess.AS2026

Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions

Abinay Reddy Naini, Jaeyeon Kim, Chao-Han Huck Yang +2

Large audio-language models (LALMs) can reason about audio, yet it remains unclear whether they can perform comparative judgments between two speech signals along emotional, enviro…

cs.LG2025

RankList -- A Listwise Preference Learning Framework for Predicting Subjective Preferences

Abinay Reddy Naini, Fernando Diaz, Carlos Busso

Preference learning has gained significant attention in tasks involving subjective human judgments, such as \emph{speech emotion recognition} (SER) and image aesthetic assessment.…

eess.AS2019

Semi-Supervised Speech Emotion Recognition with Ladder Networks

Srinivas Parthasarathy, Carlos Busso

Speech emotion recognition (SER) systems find applications in various fields such as healthcare, education, and security and defense. A major drawback of these systems is their lac…

cs.LG2025

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu +2

Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by naïvely aligning audio samples with corresponding text prompts. Consequently, this a…

cs.CV2020

The Multimodal Driver Monitoring Database: A Naturalistic Corpus to Study Driver Attention

Sumit Jha, Mohamed F. Marzban, Tiancheng Hu +3

A smart vehicle should be able to monitor the actions and behaviors of the human driver to provide critical warnings or intervene when necessary. Recent advancements in deep learni…

cs.SD2022

Unsupervised Personalization of an Emotion Recognition System: The Unique Properties of the Externalization of Valence in Speech

Kusha Sridhar, Carlos Busso

The prediction of valence from speech is an important, but challenging problem. The externalization of valence in speech has speaker-dependent cues, which contribute to performance…

eess.AS2018

Ladder Networks for Emotion Recognition: Using Unsupervised Auxiliary Tasks to Improve Predictions of Emotional Attributes

Srinivas Parthasarathy, Carlos Busso

Recognizing emotions using few attribute dimensions such as arousal, valence and dominance provides the flexibility to effectively represent complex range of emotional behaviors. C…

cs.SD2025

Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation

Karen Rosero, Eunjung Yeo, David R. Mortensen +3

We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy ad…

cs.SD2025

Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment

Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela +2

Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach th…

eess.AS2026

MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition

Luz Martinez-Lucas, Pravin Mote, Abinay Reddy Naini +2

Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions convey…

cs.LG2026

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning

Esther Sun, Bo-Hao Su, Abinay Reddy Naini +2

Speech Large Language Models (SLLMs) enable high-level emotion reasoning but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, self…

eess.AS2018

Domain Adversarial for Acoustic Emotion Recognition

Mohammed Abdelwahab, Carlos Busso

The performance of speech emotion recognition is affected by the differences in data distributions between train (source domain) and test (target domain) sets used to build and eva…

cs.AI2026

Sentipolis: Emotion-Aware Agents for Social Simulations

Chiyuan Fu, Lyuhao Chen, Yunze Xiao +3

LLM agents are increasingly used for social simulation, yet emotion is often treated as a transient cue, causing emotional amnesia and weak long-horizon continuity. We present Sent…

cs.LG2024

Versatile audio-visual learning for emotion recognition

Lucas Goncalves, Seong-Gyun Leem, Wei-Cheng Lin +2

Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only…

cs.HC2017

Speech-driven Animation with Meaningful Behaviors

Najmeh Sadoughi, Carlos Busso

Conversational agents (CAs) play an important role in human computer interaction. Creating believable movements for CAs is challenging, since the movements have to be meaningful an…

eess.AS2024

Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline

Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra +3

Voice conversion (VC) research traditionally depends on scripted or acted speech, which lacks the natural spontaneity of real-life conversations. While natural speech data is limit…

eess.AS2024

Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition

Ismail Rasim Ulgen, Zongyang Du, Carlos Busso +1

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled…

eess.AS2026

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

Aurosweta Mahapatra, Xiutian Zhao, Shreeram Suresh Chandra +7

Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and rec…

cs.HC2019

The Ambiguous World of Emotion Representation

Vidhyasaharan Sethu, Emily Mower Provost, Julien Epps +3

Artificial intelligence and machine learning systems have demonstrated huge improvements and human-level parity in a range of activities, including speech recognition, face recogni…

cs.CV2020

Estimation of Driver's Gaze Region from Head Position and Orientation using Probabilistic Confidence Regions

Sumit Jha, Carlos Busso

A smart vehicle should be able to understand human behavior and predict their actions to avoid hazardous situations. Specific traits in human behavior can be automatically predicte…

eess.AS2025

Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity

Ismail Rasim Ulgen, John H. L. Hansen, Carlos Busso +1

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning input…

eess.AS2026

Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition

Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian +4

Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency…

eess.AS2023

Mixed-EVC: Mixed Emotion Synthesis and Control in Voice Conversion

Kun Zhou, Berrak Sisman, Carlos Busso +2

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discre…

eess.AS2025

The MSP-Podcast Corpus

Carlos Busso, Reza Lotfian, Kusha Sridhar +9

The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing datab…

cs.SD2024

A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition

Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee

Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emoti…

cs.SD2025

FedMLAC: Mutual Learning Driven Heterogeneous Federated Audio Classification

Jun Bai, Rajib Rana, Di Wu +5

Federated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients without sharing raw data. However, Federate…

cs.SD2024

emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition

Thejan Rajapakshe, Rajib Rana, Sara Khalifa +3

Speech Emotion Recognition (SER) is crucial for enabling computers to understand the emotions conveyed in human communication. With recent advancements in Deep Learning (DL), the p…

eess.AS2026

Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration

Esther Sun, Abinay Reddy Naini, Carlos Busso

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguist…

eess.AS2018

Curriculum Learning for Speech Emotion Recognition from Crowdsourced Labels

Reza Lotfian, Carlos Busso

This study introduces a method to design a curriculum for machine-learning to maximize the efficiency during the training process of deep neural networks (DNNs) for speech emotion…

cs.SD2026

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

Ziyang Ma, Ruiyang Xu, Yinghao Ma +9

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Cha…

eess.AS2025

NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion

Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ulgen +4

Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted,…