198 citations · 355 across the 2 of their papers we have counts for
2 papers
cs.CV2019★ 157 cited
Context-Aware Visual Policy Network for Fine-Grained Image Captioning
Zheng-Jun Zha, Daqing Liu, Hanwang Zhang +2
With the maturity of visual detection techniques, we are more ambitious in describing visual content with open-vocabulary, fine-grained and free-form language, i.e., the task of im…
cs.CV2019★ 198 cited
Learning to Compose and Reason with Language Tree Structures for Visual Grounding
Richang Hong, Daqing Liu, Xiaoyu Mo +2
Grounding natural language in images, such as localizing "the black dog on the left of the tree", is one of the core problems in artificial intelligence, as it needs to comprehend…