4 papers
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
Nengneng Yu, Sixian Xiong, Yibo Zhao +2
Today's inference-time workloads increasingly depend on timely access to a model's internal states. We present DMI-Lib, a high-speed deep model inspector that treats internal obser…
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
Shimul Debnath, William Hart, Lori Pollock +2
Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term. Their effects are often inter…
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
Minchen Yu, Rui Yang, Chaobo Jia +9
Serverless computing has emerged as a compelling solution for cloud-based model inference. However, as modern large language models (LLMs) continue to grow in size, existing server…
Protecting Confidentiality, Privacy and Integrity in Collaborative Learning
Dong Chen, Alice Dethise, Istemi Ekin Akkus +6
A collaboration between dataset owners and model owners is needed to facilitate effective machine learning (ML) training. During this collaboration, however, dataset owners and mod…