4 papers · 1 filter
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
Dingyu Yang, Fanyong Kong, Jie Dai +5
Modern cloud servers routinely co-locate multiple latency-sensitive microservice instances to improve resource efficiency. However, the diversity of microservice behaviors, coupled…
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
Jiaang Duan, Shenglin Xu, Shiyou Qian +15
The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cl…
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
Dingyu Yang, Kangpeng Zheng, Shiyou Qian +2
Co-locating latency-critical services (LCSs) and best-effort jobs (BEJs) constitute the principal approach for enhancing resource utilization in production. Nevertheless, the co-lo…
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
Qin Hua, Dingyu Yang, Shiyou Qian +3
An effective auto-scaling framework is essential for microservices to ensure performance stability and resource efficiency under dynamic workloads. As revealed by many prior studie…