353 citations · 364 across the 2 of their papers we have counts for
1 paper · 1 filter
Tete Xiao, Mannat Singh, Eric Mintun +3
Vision transformer (ViT) models exhibit substandard optimizability. In particular, they are sensitive to the choice of optimizer (AdamW vs. SGD), optimizer hyperparameters, and tra…