4 papers
Understanding GPU Resource Interference One Level Deeper
Paul Elvinger, Foteini Strati, Natalie Enright Jerger +1
GPUs are vastly underutilized, even when running resource-intensive AI applications, as GPU kernels within each job have diverse resource profiles that may saturate some parts of a…
Arcus: SLO Management for Accelerators in the Cloud with Traffic Shaping
Jiechen Zhao, Ran Shu, Katie Lim +4
Cloud servers use accelerators for common tasks (e.g., encryption, compression, hashing) to improve CPU/GPU efficiency and overall performance. However, users' Service-level Object…
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
Jiechen Zhao, Ran Shu, Katie Lim +4
I/O devices in public clouds have integrated increasing numbers of hardware accelerators, e.g., AWS Nitro, Azure FPGA and Nvidia BlueField. However, such specialized compute (1) is…
Low-Energy Line Codes for On-Chip Networks
Beyza Dabak, Major Glenn, Jingyang Liu +5
Energy is a primary constraint in processor design, and much of that energy is consumed in on-chip communication. Communication can be intra-core (e.g., from a register file to an…