31 citations · 43 across the 5 of their papers we have counts for
1 paper · 1 filter
Alexis Marouani, Oriane Siméoni, Hervé Jégou +2
Vision Transformers have emerged as powerful, scalable and versatile representation learners. To capture both global and local features, a learnable [CLS] class token is typically…