activity
20202023
most citedCross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted Alignment

21 citations · 51 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV202321 cited

Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted Alignment

Shengqiong Wu, Hao Fei, Wei Ji +1

Unpaired cross-lingual image captioning has long suffered from irrelevancy and disfluency issues, due to the inconsistencies of the semantic scene and syntax attributes during tran…

cs.CV2022

MetaComp: Learning to Adapt for Online Depth Completion

Yang Chen, Shanshan Zhao, Wei Ji +2

Relying on deep supervised or self-supervised learning, previous methods for depth completion from paired single image and sparse depth data have achieved impressive performance in…

cs.CV2021

Rethinking the Two-Stage Framework for Grounded Situation Recognition

Meng Wei, Long Chen, Wei Ji +2

Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g., buying) and detecting all corresponding semantic roles (e.g., age…

cs.CV202116 cited

Video as Conditional Graph Hierarchy for Multi-Granular Question Answering

Junbin Xiao, Angela Yao, Zhiyuan Liu +3

Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been foc…

cs.CV202014 cited

Accurate RGB-D Salient Object Detection via Collaborative Learning

Wei Ji, Jingjing Li, Miao Zhang +2

Benefiting from the spatial cues embedded in depth images, recent progress on RGB-D saliency detection shows impressive ability on some challenge scenarios. However, there are stil…