collaborators

5 papers

cs.DC2025

Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs

Shengkun Cui, Archit Patke, Hung Nguyen +11

This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years…

cs.DC2025

CPU-Limits kill Performance: Time to rethink Resource Control

Chirag Shetty, Sarthak Chakraborty, Hubertus Franke +4

Research in compute resource management for cloud-native applications is dominated by the problem of setting optimal CPU limits -- a fundamental OS mechanism that strictly restrict…

cs.DC2025

Queue management for slo-oriented large language model serving

Archit Patke, Dhemath Reddy, Saurabh Jha +5

Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots a…

cs.AI2025

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Saurabh Jha, Rohan Arora, Yuji Watanabe +40

Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a…

cs.DC2025

Hierarchical Autoscaling for Large Language Model Serving with Chiron

Archit Patke, Dhemath Reddy, Saurabh Jha +3

Large language model (LLM) serving is becoming an increasingly important workload for cloud providers. Based on performance SLO requirements, LLM inference requests can be divided…