collaborators

6 papers

cs.CV2026

Andha-Dhun: A First Look at Audio Descriptions in Hindi

Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3

Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…

eess.AS2025

HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition

Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +4

Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based rep…

eess.AS2025

SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer

Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish +8

As video-sharing platforms have grown over the past decade, child viewership has surged, increasing the need for precise detection of harmful content like violence or explicit scen…

eess.AS2025

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +5

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind w…

eess.AS2025

Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction

Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +4

In this study, we focus on Singing Voice Mean Opinion Score (SingMOS) prediction. Previous research have shown the performance benefit with the use of state-of-the-art (SOTA) pre-t…

eess.AS2025

Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations

Girish, Mohd Mujtaba Akhtar, Orchid Chetia Phukan +5

In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and in…