8 citations · 15 across the 2 of their papers we have counts for
2 papers
cs.CV2021★ 8 cited
Can Vision Transformers Perform Convolution?
Shanda Li, Xiangning Chen, Di He +1
Several recent studies have demonstrated that attention-based networks, such as Vision Transformer (ViT), can outperform Convolutional Neural Networks (CNNs) on several computer vi…
cs.LG2021★ 7 cited
Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
Shengjie Luo, Shanda Li, Tianle Cai +6
The attention module, which is a crucial component in Transformer, cannot scale efficiently to long sequences due to its quadratic complexity. Many works focus on approximating the…