most citedCrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

85 citations · 103 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV202185 cited

CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

Wenxiao Wang, Lu Yao, Long Chen +4

Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among…

cs.LG202113 cited

Attacking Adversarial Attacks as A Defense

Boxi Wu, Heng Pan, Li Shen +6

It is well known that adversarial attacks can fool deep neural networks with imperceptible perturbations. Although adversarial training significantly improves model robustness, fai…

cs.CV20215 cited

Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Long Chen, Zhihong Jiang, Jun Xiao +1

Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulat…

cs.CV2021

A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric

Yitian Yuan, Xiaohan Lan, Xin Wang +3

Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has recei…

cs.CV2020

Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding

Long Chen, Wenbo Ma, Jun Xiao +2

The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to…