collaborators

5 papers

cs.DC2025

Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs

Shengkun Cui, Archit Patke, Hung Nguyen +11

This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years…

cs.DC2025

INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network

Archit Patke, Christian Pinto, Saurabh Jha +3

Hardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in…

cs.DC2025

Queue management for slo-oriented large language model serving

Archit Patke, Dhemath Reddy, Saurabh Jha +5

Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots a…

cs.DC2025

Hierarchical Autoscaling for Large Language Model Serving with Chiron

Archit Patke, Dhemath Reddy, Saurabh Jha +3

Large language model (LLM) serving is becoming an increasingly important workload for cloud providers. Based on performance SLO requirements, LLM inference requests can be divided…

cs.DC2024

Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Haoran Qiu, Weichao Mao, Archit Patke +7

Large language models (LLMs) have been driving a new wave of interactive AI applications across numerous domains. However, efficiently serving LLM inference requests is challenging…