3 citations · 3 across the 5 of their papers we have counts for
7 papers · 1 filter
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
Archit Patke, Christian Pinto, Saurabh Jha +3
Hardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in…
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
Shengkun Cui, Archit Patke, Hung Nguyen +11
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years…
Hierarchical Autoscaling for Large Language Model Serving with Chiron
Archit Patke, Dhemath Reddy, Saurabh Jha +3
Large language model (LLM) serving is becoming an increasingly important workload for cloud providers. Based on performance SLO requirements, LLM inference requests can be divided…
Queue management for slo-oriented large language model serving
Archit Patke, Dhemath Reddy, Saurabh Jha +5
Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots a…
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Haoran Qiu, Weichao Mao, Archit Patke +7
Large language models (LLMs) have been driving a new wave of interactive AI applications across numerous domains. However, efficiently serving LLM inference requests is challenging…
Application-aware Congestion Mitigation for High-Performance Computing Systems
Archit Patke, Saurabh Jha, Haoran Qiu +5
High-performance computing (HPC) systems frequently experience congestion leading to significant application performance variation. However, the impact of congestion on application…