486 citations · 508 across the 8 of their papers we have counts for
12 papers · 1 filter
Deconfounded Video Moment Retrieval with Causal Intervention
Xun Yang, Fuli Feng, Wei Ji +2
We tackle the task of video moment retrieval (VMR), which aims to localize a specific moment in a video according to a textual query. Existing methods primarily model the matching…
Feature Pyramid Transformer
Dong Zhang, Hanwang Zhang, Jinhui Tang +3
Feature interactions across space and scales underpin modern visual recognition systems because they introduce beneficial visual contexts. Conventionally, spatial contexts are pass…
Learning to Discretely Compose Reasoning Module Networks for Video Captioning
Ganchao Tan, Daqing Liu, Meng Wang +1
Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the…
Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval
Xun Yang, Jianfeng Dong, Yixin Cao +3
The rapid growth of user-generated videos on the Internet has intensified the need for text-based video retrieval systems. Traditional methods mainly favor the concept-based paradi…
Memory-Augmented Relation Network for Few-Shot Learning
Jun He, Richang Hong, Xueliang Liu +3
Metric-based few-shot learning methods concentrate on learning transferable feature embedding that generalizes well from seen categories to unseen categories under the supervision…
Iterative Context-Aware Graph Inference for Visual Dialog
Dan Guo, Hui Wang, Hanwang Zhang +2
Visual dialog is a challenging task that requires the comprehension of the semantic dependencies among implicit visual and textual contexts. This task can refer to the relation inf…