Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Enhancing Audio Captioning with Auxiliary AudioSet Semantics
Shubham Gupta, Adarsh Arigala, Sri Rama Murty Kodukula
Automatic Audio Captioning (AAC) seeks to generate natural language descriptions of complex acoustic scenes, bridging auditory perception and language understanding. However, word-…
eess.AS2026
Shared Representation Learning for Reference-Guided Targeted Sound Detection
Shubham Gupta, Adarsh Arigala, B. R. Dilleswari +1
Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Soun…
eess.AS2024
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
Shubham Gupta, Mirco Ravanelli, Pascal Germain +1
In this paper, we propose Phoneme Discretized Saliency Maps (PDSM), a discretization algorithm for saliency maps that takes advantage of phoneme boundaries for explainable detectio…