1 paper
Naren Dhyani, Jianqiao Mo, Minsu Cho +4
The Vision Transformer (ViT) architecture has emerged as the backbone of choice for state-of-the-art deep models for computer vision applications. However, ViTs are ill-suited for…