activity
20242026
collaborators

15 papers

eess.AS2026

Towards Detecting Neural Audio Codec Synthesized Heart Sounds

Girish, Orchid Chetia Phukan, Mohd Mujtaba Akhtar +3

In this paper, we introduce Synthetic Heart Sound Detection (SHAC), a task aimed at identifying phonocardiograms (PCGs) synthesized using neural audio codecs (NACs). To facilitate…

eess.AS2025

Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?

Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +4

Recent benchmarks evaluating pre-trained models (PTMs) for cross-corpus speech emotion recognition (SER) have overlooked PTM pre-trained for paralinguistic speech processing (PSP),…

eess.AS2025

HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition

Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +4

Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based rep…

eess.AS2025

SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer

Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +8

As video-sharing platforms have grown over the past decade, child viewership has surged, increasing the need for precise detection of harmful content like violence or explicit scen…

eess.AS2025

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +5

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind w…

eess.AS2025

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

Mohd Mujtaba Akhtar, Orchid Chetia Phukan, Girish +5

In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize…