activity
20172022
most citedAn End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning

11 citations · 18 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20225 cited

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian +6

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language…

cs.CV2021

Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language Inference

Juncheng Li, Siliang Tang, Linchao Zhu +5

Video-and-Language Inference is a recently proposed task for joint video-and-language understanding. This new task requires a model to draw inference on whether a natural language…

cs.CV20202 cited

Grounded and Controllable Image Completion by Incorporating Lexical Semantics

Shengyu Zhang, Tan Jiang, Qinghao Huang +7

In this paper, we present an approach, namely Lexical Semantic Image Completion (LSIC), that may have potential applications in art, design, and heritage conservation, among severa…

cs.CV2018

Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions

Ke Ning, Linchao Zhu, Ming Cai +3

We propose a novel attentive sequence to sequence translator (ASST) for clip localization in videos by natural language descriptions. We make two contributions. First, we propose a…

cs.CV201711 cited

An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning

Fan Wu, Zhongwen Xu, Yi Yang

We propose an end-to-end approach to the natural language object retrieval task, which localizes an object within an image according to a natural language description, i.e., referr…