Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Wentao Liu, Yuhao Hu, Ruiting Zhou +2
Mixture-of-Experts (MoE) has become a dominant architecture in large language models (LLMs) due to its ability to scale model capacity via sparse expert activation. Meanwhile, serv…
cs.DC2021
Online Service Placement and Request Scheduling in MEC Networks
Lina Su, Ne Wang, Ruiting Zhou +1
Mobile edge computing (MEC) emerges as a promising solution for servicing delay-sensitive tasks at the edge network. A body of recent literature started to focus on cost-efficient…