1 paper · 1 filter
Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz +2
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In…