2 papers
cs.DC2026
Making MoE-based LLM Inference Resilient with Tarragon
Songyu Zhang, Aaron Tam, Myungjin Lee +2
Mixture-of-Experts (MoE) models are increasingly used to serve LLMs at scale, but failures become common as deployment scale grows. Existing systems exhibit poor failure resilience…
cs.DC2024
LIFL: A Lightweight, Event-driven Serverless Platform for Federated Learning
Shixiong Qi, K. K. Ramakrishnan, Myungjin Lee
Federated Learning (FL) typically involves a large-scale, distributed system with individual user devices/servers training models locally and then aggregating their model updates o…