2 papers
cs.CL2026
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
cs.DC2025
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
Heng Xu, Zhiwei Yu, Chengze Du +5
Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to…