2 papers
cs.DC2025
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
Shengkun Cui, Archit Patke, Hung Nguyen +11
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years…
cs.CR2024
Security Testbed for Preempting Attacks against Supercomputing Infrastructure
Phuong Cao, Zbigniew Kalbarczyk, Ravishankar Iyer
Securing HPC has a unique threat model. Untrusted, malicious code exploiting the concentrated computing power may exert an outsized impact on the shared, open-networked environment…