2 papers
cs.LG2026
MobileMoE: Scaling On-Device Mixture of Experts
Yanbei Chen, Hanxian Huang, Ernie Chang +5
Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remai…
cs.LG2026
ExecuTorch -- A Unified PyTorch Solution to Run AI Models On-Device
Mergen Nachin, Digant Desai, Sicheng Stephen Jia +36
Local execution of AI on edge devices is important for low latency and offline operation. However, deploying models on diverse hardware remains fragmented, often requiring model co…