14 papers
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
Girish, Mohd Mujtaba Akhtar, Farhan Sheth +1
In this work, we address the problem of finegrained traceback of emotional and manipulation characteristics from synthetically manipulated speech. We hypothesize that combining sem…
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +6
This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse l…
Are Multimodal Foundation Models All That Is Needed for Emofake Detection?
Mohd Mujtaba Akhtar, Girish, Orchid Chetia Phukan +5
In this work, we investigate multimodal foundation models (MFMs) for EmoFake detection (EFD) and hypothesize that they will outperform audio foundation models (AFMs). MFMs due to t…
Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +4
Recent benchmarks evaluating pre-trained models (PTMs) for cross-corpus speech emotion recognition (SER) have overlooked PTM pre-trained for paralinguistic speech processing (PSP),…
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +1
In this work, we address EmoFake Detection (EFD). We hypothesize that multilingual speech foundation models (SFMs) will be particularly effective for EFD due to their pre-training…
Towards Neural Audio Codec Source Parsing
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +2
A new class of audio deepfakes-codecfakes (CFs)-has recently caught attention, synthesized by Audio Language Models that leverage neural audio codecs (NACs) in the backend. In resp…