4 citations · 4 across the 4 of their papers we have counts for
1 paper · 1 filter
Yukun Jiang, Hai Huang, Mingjie Li +3
By introducing routers to selectively activate experts in Transformer layers, the mixture-of-experts (MoE) architecture significantly reduces computational costs in large language…