1 paper · 1 filter
Xiaocheng Zou, Tiancheng Zheng, Xiaolin Xu +1
Mixture-of-Experts (MoE) has become a prevalent paradigm for scaling Vision Transformers efficiently. To ensure computational scalability and prevent expert overload, Vision MoE ar…