1 paper
Jiale Dong, Wenqi Lou, Zhendong Zheng +4
Compared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computatio…