1 paper
Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10
In this work, we introduce Instella-MoE, a fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters per token, trained…