activity
20232025
most citedLeveraging Foundation models for Unsupervised Audio-Visual Segmentation

3 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.IR2025

NEAR: A Nested Embedding Approach to Efficient Product Retrieval and Ranking

Shenbin Qian, Diptesh Kanojia, Samarth Agrawal +4

E-commerce information retrieval (IR) systems struggle to simultaneously achieve high accuracy in interpreting complex user queries and maintain efficient processing of vast produc…

cs.CV20241 cited

Unsupervised Audio-Visual Segmentation with Modality Alignment

Swapnil Bhosale, Haosen Yang, Diptesh Kanojia +2

Audio-Visual Segmentation (AVS) aims to identify, at the pixel level, the object in a visual scene that produces a given sound. Current AVS methods rely on costly fine-grained anno…

cs.CL20231 cited

Sarcasm in Sight and Sound: Benchmarking and Expansion to Improve Multimodal Sarcasm Detection

Swapnil Bhosale, Abhra Chaudhuri, Alex Lee Robert Williams +5

The introduction of the MUStARD dataset, and its emotion recognition extension MUStARD++, have identified sarcasm to be a multi-modal phenomenon -- expressed not only in natural la…

cs.CV20233 cited

Leveraging Foundation models for Unsupervised Audio-Visual Segmentation

Swapnil Bhosale, Haosen Yang, Diptesh Kanojia +1

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask…

cs.SD2023

DiffSED: Sound Event Detection with Denoising Diffusion

Swapnil Bhosale, Sauradip Nag, Diptesh Kanojia +2

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the spl…