activity
20172022
most citedVideo Question Answering via Attribute-Augmented Attention Network Learning

103 citations · 232 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20228 cited

Respecting Transfer Gap in Knowledge Distillation

Yulei Niu, Long Chen, Chang Zhou +1

Knowledge distillation (KD) is essentially a process of transferring a teacher model's behavior, e.g., network response, to a student model. The network response serves as addition…

cs.CV20221 cited

A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach

Xiaohan Lan, Yitian Yuan, Xin Wang +4

Temporal Sentence Grounding in Videos (TSGV), which aims to ground a natural language sentence in an untrimmed video, has drawn widespread attention over the past few years. Howeve…

cs.CV202118 cited

Video Relation Detection via Tracklet based Visual Transformer

Kaifeng Gao, Long Chen, Yifeng Huang +1

Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet…

cs.CV202185 cited

CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

Wenxiao Wang, Lu Yao, Long Chen +4

Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among…

cs.CV20215 cited

Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Long Chen, Zhihong Jiang, Jun Xiao +1

Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulat…

cs.CV2021

A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric

Yitian Yuan, Xiaohan Lan, Xin Wang +3

Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has recei…