1 paper
Yinjie Chen, Zipeng Yan, Chong Zhou +2
Vision Transformers (ViTs) have emerged as the dominant architecture for visual processing tasks, demonstrating excellent scalability with increased training data and model size. H…