1 citations · 1 across the 3 of their papers we have counts for
5 papers
Knowledge-Guided Object Discovery with Acquired Deep Impressions
Jinyang Yuan, Bin Li, Xiangyang Xue
We present a framework called Acquired Deep Impressions (ADI) which continuously learns knowledge of objects as "impressions" for compositional scene understanding. In this framewo…
A Generic Object Re-identification System for Short Videos
Tairu Qiu, Guanxian Chen, Zhongang Qi +3
Short video applications like TikTok and Kwai have been a great hit recently. In order to meet the increasing demands and take full advantage of visual information in short videos,…
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Weijie Su, Xizhou Zhu, Yue Cao +4
We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Trans…
Question Guided Modular Routing Networks for Visual Question Answering
Yanze Wu, Qiang Sun, Jianqi Ma +4
This paper studies the task of Visual Question Answering (VQA), which is topical in Multimedia community recently. Particularly, we explore two critical research problems existed i…
Spatial Mixture Models with Learnable Deep Priors for Perceptual Grouping
Jinyang Yuan, Bin Li, Xiangyang Xue
Humans perceive the seemingly chaotic world in a structured and compositional way with the prerequisite of being able to segregate conceptual entities from the complex visual scene…