118 citations · 118 across the 2 of their papers we have counts for
3 papers
Image Captioners Are Scalable Vision Learners Too
Michael Tschannen, Manoj Kumar, Andreas Steiner +3
Contrastive pretraining on image-text pairs from the web is one of the most popular large-scale pretraining strategies for vision backbones, especially in the context of large mult…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
Dual PatchNorm
Manoj Kumar, Mostafa Dehghani, Neil Houlsby
We propose Dual PatchNorm: two Layer Normalization layers (LayerNorms), before and after the patch embedding layer in Vision Transformers. We demonstrate that Dual PatchNorm outper…