16 citations · 30 across the 3 of their papers we have counts for
3 papers
cs.CV2021
Rethinking the Two-Stage Framework for Grounded Situation Recognition
Meng Wei, Long Chen, Wei Ji +2
Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g., buying) and detecting all corresponding semantic roles (e.g., age…
cs.CV2021★ 16 cited
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering
Junbin Xiao, Angela Yao, Zhiyuan Liu +3
Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been foc…
cs.CV2020★ 14 cited
Accurate RGB-D Salient Object Detection via Collaborative Learning
Wei Ji, Jingjing Li, Miao Zhang +2
Benefiting from the spatial cues embedded in depth images, recent progress on RGB-D saliency detection shows impressive ability on some challenge scenarios. However, there are stil…