20 citations · 34 across the 13 of their papers we have counts for
1 paper · 1 filter
Naren Dhyani, Jianqiao Mo, Minsu Cho +4
The Vision Transformer (ViT) architecture has emerged as the backbone of choice for state-of-the-art deep models for computer vision applications. However, ViTs are ill-suited for…