794 citations · 2.4k across the 85 of their papers we have counts for
9 papers · 1 filter
Visual Question Generation as Dual Task of Visual Question Answering
Yikang Li, Nan Duan, Bolei Zhou +3
Recently visual question answering (VQA) and visual question generation (VQG) are two trending topics in the computer vision, which have been explored separately. In this work, we…
Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism
Qi Chu, Wanli Ouyang, Hongsheng Li +3
In this paper, we propose a CNN-based framework for online MOT. This framework utilizes the merits of single object trackers in adapting appearance models and searching for target…
Learning Feature Pyramids for Human Pose Estimation
Wei Yang, Shuang Li, Wanli Ouyang +2
Articulated human pose estimation is a fundamental yet challenging task in computer vision. The difficulty is particularly pronounced in scale variations of human body parts when c…
Scene Graph Generation from Objects, Phrases and Region Captions
Yikang Li, Wanli Ouyang, Bolei Zhou +2
Object detection, scene graph generation and region captioning, which are three scene understanding tasks at different semantic levels, are tied together: scene graphs are generate…
Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
Zhe Wang, Hongsheng Li, Wanli Ouyang +1
Scene labeling is a challenging classification problem where each input image requires a pixel-level prediction map. Recently, deep-learning-based methods have shown their effectiv…
Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
Feng Zhu, Hongsheng Li, Wanli Ouyang +2
Multi-label image classification is a fundamental but challenging task in computer vision. Great progress has been achieved by exploiting semantic relations between labels in recen…