1 paper
Sugyeong Eo, Jungjun Lee, Chanjun Park +1
A sparse Mixture-of-Experts (MoE) architecture has emerged as a highly scalable solution by conditionally activating sub-modules without a proportional increase in computational co…