most citedLearning Cross-modal Context Graph for Visual Grounding

8 citations · 10 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2021

Single Image 3D Object Estimation with Primitive Graph Networks

Qian He, Desen Zhou, Bo Wan +1

Reconstructing 3D object from a single image (RGB or depth) is a fundamental problem in visual scene understanding and yet remains challenging due to its ill-posed nature and compl…

cs.CV2021

Bipartite Graph Network with Adaptive Message Passing for Unbiased Scene Graph Generation

Rongjie Li, Songyang Zhang, Bo Wan +1

Scene graph generation is an important visual understanding task with a broad range of vision applications. Despite recent tremendous progress, it remains challenging due to the in…

cs.CV20212 cited

Relation-aware Instance Refinement for Weakly Supervised Visual Grounding

Yongfei Liu, Bo Wan, Lin Ma +1

Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and sc…

cs.CV20198 cited

Learning Cross-modal Context Graph for Visual Grounding

Yongfei Liu, Bo Wan, Xiaodan Zhu +1

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding ent…

cs.CV2019

Pose-aware Multi-level Feature Network for Human Object Interaction Detection

Bo Wan, Desen Zhou, Yongfei Liu +2

Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large vari…