21 papers
Towards Detecting Neural Audio Codec Synthesized Heart Sounds
Girish, Orchid Chetia Phukan, Mohd Mujtaba Akhtar +3
In this paper, we introduce Synthetic Heart Sound Detection (SHAC), a task aimed at identifying phonocardiograms (PCGs) synthesized using neural audio codecs (NACs). To facilitate…
Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar +1
Codecfakes (CFs) are a type of speech deepfakes generated through Audio Language Models (ALMs), with Neural Audio Codecs (NACs) forming the core mechanism for speech encoding and g…
DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition
Bikrant Bikram Pratap Maurya, Nitin Choudhury, Daksh Agarwal +1
Acoustic side-channel attacks (ASCA) on keyboards pose a significant security risk, as keystrokes can be inferred from typing acoustics, revealing sensitive information. Prior ASCA…
RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System
Nitin Choudhury, Nikhil Kumar, Aditya Kumar Sinha +5
Wide exploration on robocall surveillance research is hindered due to limited access to public datasets, due to privacy concerns. In this work, we first curate Robo-SAr, a syntheti…
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
Girish, Mohd Mujtaba Akhtar, Orchid Chetia Phukan +1
The rapid advancement of Audio Large Language Models (ALMs), driven by Neural Audio Codecs (NACs), has led to the emergence of highly realistic speech deepfakes, commonly referred…
FOCA: Multimodal Malware Classification via Hyperbolic Cross-Attention
Nitin Choudhury, Bikrant Bikram Pratap Maurya, Orchid Chetia Phukan +1
In this work, we introduce FOCA, a novel multimodal framework for malware classification that jointly leverages audio and visual modalities. Unlike conventional Euclidean-based fus…