5 papers
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
Shaoke Xi, ChonLam Lao, Boyi Jia +11
Large language model (LLM) training today runs on clusters spanning thousands of GPUs. While this scale enables rapid model advances, developing, debugging, and performance-tuning…
Accelerating Compound LLM Training Workloads with Maestro
Xiulong Yuan, Hongqing Chen, Jiaxuan Peng +16
Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differin…
TrainMover: An Interruption-Resilient Runtime for ML Training
ChonLam Lao, Jiaqi Gao, Jiamin Cao +13
Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…
Cora: Accelerating Stateful Network Applications with SmartNICs
Shaoke Xi, Jiaqi Gao, Mengqi Liu +7
With the growing performance requirements on networked applications, there is a new trend of offloading stateful network applications to SmartNICs to improve performance and reduce…
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
ChonLam Lao, Jiaqi Gao, Ganesh Ananthanarayanan +2
Traditional ML inference is evolving toward modeless inference, which abstracts the complexity of model selection from users, allowing the system to automatically choose the most a…