1 paper
Junghwan Lim, Joon Son Chung, Sungmin Lee +24
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 ro…