2 papers
cs.AI2024
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations
Fucheng Jia, Shiqi Jiang, Ting Cao +9
Web is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterog…
cs.AI2024
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
Rui Kong, Yuanchun Li, Qingtian Feng +5
Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) with conditionally-activated parallel experts. However, serving MoE models on me…