12 citations · 14 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 2 cited
Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning
Weixuan Sun, Jiayi Zhang, Jianyuan Wang +6
Self-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the hel…
cs.CV2023★ 12 cited
Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Jianyuan Wang +8
We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame…