44 citations · 63 across the 5 of their papers we have counts for
5 papers
You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
Chaorui Deng, Qi Wu, Guanghui Xu +4
Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful tech…
Neighbourhood Watch: Referring Expression Comprehension via Language-guided Graph Attention Networks
Peng Wang, Qi Wu, Jiewei Cao +3
The task in referring expression comprehension is to localise the object instance in an image described by a referring expression phrased in natural language. As a language-to-visi…
The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
Peng Wang, Qi Wu, Chunhua Shen +1
One of the most intriguing features of the Visual Question Answering (VQA) challenge is the unpredictability of the questions. Extracting the information required to answer them de…
Multi-Label Image Classification with Regional Latent Semantic Dependencies
Junjie Zhang, Qi Wu, Chunhua Shen +2
Deep convolution neural networks (CNN) have demonstrated advanced performance on single-label image classification, and various progress also have been made to apply CNN methods on…
Visual Question Answering: A Survey of Methods and Datasets
Qi Wu, Damien Teney, Peng Wang +3
Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities. Given an…