2 papers
cs.CV2025
Audio-Visual Instance Segmentation
Ruohao Guo, Xianghua Ying, Yaru Chen +11
In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding obj…
cs.MM2024
Open-Vocabulary Audio-Visual Semantic Segmentation
Ruohao Guo, Liao Qu, Dantong Niu +5
Audio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption a…