1 citations · 1 across the 1 of their papers we have counts for
4 papers · 1 filter
Length-Controllable Image Captioning
Chaorui Deng, Ning Ding, Mingkui Tan +1
The last decade has witnessed remarkable progress in the image captioning task; however, most existing methods cannot control their captions, \emph{e.g.}, choosing to describe the…
Referring Expression Comprehension: A Survey of Methods and Datasets
Yanyuan Qiao, Chaorui Deng, Qi Wu
Referring expression comprehension (REC) aims to localize a target object in an image described by a referring expression phrased in natural language. Different from the object det…
Deep High-Resolution Representation Learning for Visual Recognition
Jingdong Wang, Ke Sun, Tianheng Cheng +9
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-…
You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
Chaorui Deng, Qi Wu, Guanghui Xu +4
Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful tech…