1 paper
Harsh Rangwani, Pradipto Mondal, Mayank Mishra +2
Vision Transformer (ViT) has emerged as a prominent architecture for various computer vision tasks. In ViT, we divide the input image into patch tokens and process them through a s…