1 paper
Hongyu Wang, Jiayu Xu, Ruiping Wang +5
Large multimodal Mixture-of-Experts (MoEs) effectively scale the model size to boost performance while maintaining fixed active parameters. However, previous works primarily utiliz…