1 paper · 1 filter
Youngseog Chung, Dhruv Malik, Jeff Schneider +2
The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small e…