3 papers
cs.DC2025
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Wentao Liu, Yuhao Hu, Ruiting Zhou +2
Mixture-of-Experts (MoE) has become a dominant architecture in large language models (LLMs) due to its ability to scale model capacity via sparse expert activation. Meanwhile, serv…
cs.LG2025
AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
Mengyang Liu, Chenyu Lu, Haodong Tian +7
On-device machine learning (ML) has become a fundamental component of emerging mobile applications. Adaptive model deployment delivers efficient inference for heterogeneous device…
cs.NI2024
Diagnosing and Repairing Distributed Routing Configurations Using Selective Symbolic Simulation
Rulan Yang, Gao Han, Hanyang Shao +9
Although substantial progress has been made in automatically verifying whether distributed routing configurations conform to certain requirements, diagnosing and repairing configur…