56 citations · 139 across the 7 of their papers we have counts for
8 papers · 1 filter
Group Generalized Mean Pooling for Vision Transformer
Byungsoo Ko, Han-Gyu Kim, Byeongho Heo +4
Vision Transformer (ViT) extracts the final representation from either class token or an average of all patch tokens, following the architecture of Transformer in Natural Language…
An Extendable, Efficient and Effective Transformer-based Object Detector
Hwanjun Song, Deqing Sun, Sanghyuk Chun +5
Transformers have been widely used in numerous vision problems especially for visual recognition and detection. Detection transformers are the first fully end-to-end learning syste…
Learning Features with Parameter-Free Layers
Dongyoon Han, YoungJoon Yoo, Beomyoung Kim +1
Trainable layers such as convolutional building blocks are the standard network design choices by learning parameters to capture the global context through successive spatial opera…
Rethinking Spatial Dimensions of Vision Transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han +3
Vision Transformer (ViT) extends the application range of transformers from language processing to computer vision tasks as being an alternative architecture against the existing c…
Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
Sangdoo Yun, Seong Joon Oh, Byeongho Heo +3
ImageNet has been arguably the most popular image classification benchmark, but it is also the one with a significant level of label noise. Recent studies have shown that many samp…
VideoMix: Rethinking Data Augmentation for Video Classification
Sangdoo Yun, Seong Joon Oh, Byeongho Heo +2
State-of-the-art video action classifiers often suffer from overfitting. They tend to be biased towards specific objects and scene cues, rather than the foreground action content,…