1 paper
Chaoxiang Cai, Longrong Yang, Minghe Weng +3
The mixture-of-experts (MoE) architecture, which replaces dense networks with sparse ones, has attracted significant attention in large vision-language models (LVLMs) for achieving…