1 paper · 1 filter
Anantha Padmanaban Krishna Kumar
Although scaling laws and many empirical results suggest that increasing the size of Vision Transformers often improves performance, model accuracy and training behavior are not al…