18 citations · 30 across the 6 of their papers we have counts for
1 paper · 1 filter
Naren Dhyani, Jianqiao Mo, Minsu Cho +4
The Vision Transformer (ViT) architecture has emerged as the backbone of choice for state-of-the-art deep models for computer vision applications. However, ViTs are ill-suited for…