Sung Jun Cheon, Jaekyung Cho, Seongho Choi +57
We introduce A.X K1, a 519B-parameter Mixture-of-Experts (MoE) language model trained from scratch. Our design leverages scaling laws to optimize training configurations and vocabu…