1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Leyla Naz Candogan, Arshia Afzal, Pol Puigdemont +1
Though Vision Transformers (ViTs) have become the dominant backbone in many computer vision tasks, due to permutation equivariance, their attention mechanism lacks explicit spatial…