1 paper · 1 filter
Foozhan Ataiefard, Walid Ahmed, Habib Hajimolahoseini +7
Vision transformers are known to be more computationally and data-intensive than CNN models. These transformer models such as ViT, require all the input image tokens to learn the r…