13 papers
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
Gabin Schieffer, Ruimin Shi, Jie Ren +1
Advances in GPU compute throughput and memory capacity brings significant opportunities to a wide range of workloads. However, efficiently utilizing these resources remains challen…
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
Ruimin Shi, Gabin Schieffer, Pei-Hung Lin +3
ARM SVE and RISC-V RVV are emerging vector architectures in high-end processors that support vectorization of flexible vector length. In this work, we leverage an important workloa…
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
Jacob Wahlgren, Gabin Schieffer, Ruimin Shi +4
Discrete GPUs are a cornerstone of HPC and data center systems, requiring management of separate CPU and GPU memory spaces. Unified Virtual Memory (UVM) has been proposed to ease t…
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
Gabin Schieffer, Jacob Wahlgren, Ruimin Shi +4
The ever-increasing compute performance of GPU accelerators drives up the need for efficient data movements within HPC applications to sustain performance. Proposed as a solution t…
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Ruimin Shi, Gabin Schieffer, Maya Gokhale +3
Vector architectures are essential for boosting computing throughput. ARM provides SVE as the next-generation length-agnostic vector extension beyond traditional fixed-length SIMD.…
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
Gabin Schieffer, Stefano Markidis, Ivy Peng
Quantum computer simulators are an indispensable tool for prototyping quantum algorithms and verifying the functioning of existing quantum computer hardware. The current largest qu…