85 citations · 103 across the 3 of their papers we have counts for
5 papers
CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang, Lu Yao, Long Chen +4
Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among…
Attacking Adversarial Attacks as A Defense
Boxi Wu, Heng Pan, Li Shen +6
It is well known that adversarial attacks can fool deep neural networks with imperceptible perturbations. Although adversarial training significantly improves model robustness, fai…
Human-like Controllable Image Captioning with Verb-specific Semantic Roles
Long Chen, Zhihong Jiang, Jun Xiao +1
Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulat…
A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric
Yitian Yuan, Xiaohan Lan, Xin Wang +3
Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has recei…
Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
Long Chen, Wenbo Ma, Jun Xiao +2
The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to…