activity
20152021
most citedCrowded Scene Analysis: A Survey

486 citations · 508 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV20211 cited

Deconfounded Video Moment Retrieval with Causal Intervention

Xun Yang, Fuli Feng, Wei Ji +2

We tackle the task of video moment retrieval (VMR), which aims to localize a specific moment in a video according to a textual query. Existing methods primarily model the matching…

cs.CV20207 cited

Feature Pyramid Transformer

Dong Zhang, Hanwang Zhang, Jinhui Tang +3

Feature interactions across space and scales underpin modern visual recognition systems because they introduce beneficial visual contexts. Conventionally, spatial contexts are pass…

cs.CV20206 cited

Learning to Discretely Compose Reasoning Module Networks for Video Captioning

Ganchao Tan, Daqing Liu, Meng Wang +1

Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the…

cs.CV20202 cited

Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval

Xun Yang, Jianfeng Dong, Yixin Cao +3

The rapid growth of user-generated videos on the Internet has intensified the need for text-based video retrieval systems. Traditional methods mainly favor the concept-based paradi…

cs.CV20201 cited

Memory-Augmented Relation Network for Few-Shot Learning

Jun He, Richang Hong, Xueliang Liu +3

Metric-based few-shot learning methods concentrate on learning transferable feature embedding that generalizes well from seen categories to unseen categories under the supervision…

cs.CV20203 cited

Iterative Context-Aware Graph Inference for Visual Dialog

Dan Guo, Hui Wang, Hanwang Zhang +2

Visual dialog is a challenging task that requires the comprehension of the semantic dependencies among implicit visual and textual contexts. This task can refer to the relation inf…