8 citations · 10 across the 3 of their papers we have counts for
4 papers
GEM: A General Evaluation Benchmark for Multimodal Tasks
Lin Su, Nan Duan, Edward Cui +7
In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus o…
Relation-aware Instance Refinement for Weakly Supervised Visual Grounding
Yongfei Liu, Bo Wan, Lin Ma +1
Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and sc…
Learning Cross-modal Context Graph for Visual Grounding
Yongfei Liu, Bo Wan, Xiaodan Zhu +1
Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding ent…
Pose-aware Multi-level Feature Network for Human Object Interaction Detection
Bo Wan, Desen Zhou, Yongfei Liu +2
Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large vari…