2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 2 cited
Mutual Information Regularization for Weakly-supervised RGB-D Salient Object Detection
Aixuan Li, Yuxin Mao, Jing Zhang +1
In this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective mult…
cs.CV2023
Fine-grained Audible Video Description
Xuyang Shen, Dong Li, Jinxing Zhou +9
We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audibl…