2 papers
cs.LG2026
Load Testing for Machine Learning Model Serving Systems at Scale
Amr S. Abdelfattah, Nakul Tirumalai, Indu Mohanan +4
Machine learning (ML) model serving has become a dominant consumer of GPU infrastructure, yet capacity planning in these systems remains largely ad hoc. Under-provisioning leads to…
cs.LG2026
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6
Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportional per-token compute, enabling…