103 citations · 232 across the 7 of their papers we have counts for
10 papers · 1 filter
Respecting Transfer Gap in Knowledge Distillation
Yulei Niu, Long Chen, Chang Zhou +1
Knowledge distillation (KD) is essentially a process of transferring a teacher model's behavior, e.g., network response, to a student model. The network response serves as addition…
A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach
Xiaohan Lan, Yitian Yuan, Xin Wang +4
Temporal Sentence Grounding in Videos (TSGV), which aims to ground a natural language sentence in an untrimmed video, has drawn widespread attention over the past few years. Howeve…
Video Relation Detection via Tracklet based Visual Transformer
Kaifeng Gao, Long Chen, Yifeng Huang +1
Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet…
CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang, Lu Yao, Long Chen +4
Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among…
Human-like Controllable Image Captioning with Verb-specific Semantic Roles
Long Chen, Zhihong Jiang, Jun Xiao +1
Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulat…
A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric
Yitian Yuan, Xiaohan Lan, Xin Wang +3
Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has recei…