3 papers
cs.DC2025
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
Shashwat Jaiswal, Kunal Jain, Yogesh Simmhan +9
Global cloud service providers handle inference workloads for Large Language Models (LLMs) that span latency-sensitive (e.g., chatbots) and insensitive (e.g., report writing) tasks…
cs.DC2025
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
Kunal Jain, Anjaly Parayil, Ankur Mallick +10
Large Language Model (LLM) workloads have distinct prefill and decode phases with different compute and memory requirements which should ideally be accounted for when scheduling in…
cs.LG2024
Ensuring Fair LLM Serving Amid Diverse Applications
Redwan Ibne Seraj Khan, Kunal Jain, Haiying Shen +12
In a multi-tenant large language model (LLM) serving platform hosting diverse applications, some users may submit an excessive number of requests, causing the service to become una…