1 citations
1 paper
Ahmed Aldahdooh, Wassim Hamidouche, Olivier Deforges
The major part of the vanilla vision transformer (ViT) is the attention block that brings the power of mimicking the global context of the input image. For better performance, ViT…