1 paper · 1 filter
Guoyang Xia, Yifeng Ding, Fengfa Li +4
Mixture of Experts (MoE) architectures have become a key approach for scaling large language models, with growing interest in extending them to multimodal tasks. Existing methods t…