10 papers · 1 filter
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
Aurosweta Mahapatra, Xiutian Zhao, Shreeram Suresh Chandra +7
Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and rec…
Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions
Abinay Reddy Naini, Jaeyeon Kim, Chao-Han Huck Yang +2
Large audio-language models (LALMs) can reason about audio, yet it remains unclear whether they can perform comparative judgments between two speech signals along emotional, enviro…
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
Esther Sun, Abinay Reddy Naini, Carlos Busso
Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguist…
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ulgen +4
Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted,…
The MSP-Podcast Corpus
Carlos Busso, Reza Lotfian, Kusha Sridhar +9
The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing datab…
Can Emotion Fool Anti-spoofing?
Aurosweta Mahapatra, Ismail Rasim Ulgen, Abinay Reddy Naini +2
Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness…