4 papers
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +5
The adaptation of foundation models has significantly advanced environmental audio deepfake detection (EADD), a rapidly growing area of research. These models are typically fine-tu…
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
Orchid Chetia Phukan, Swarup Ranjan Behera, Shubham Singh +5
In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. Whil…
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +5
In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across mul…
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
Orchid Chetia Phukan, Sarthak Jain, Swarup Ranjan Behera +3
In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detec…