activity
20242026
collaborators

13 papers

cs.DC2026

Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading

Gabin Schieffer, Ruimin Shi, Jie Ren +1

Advances in GPU compute throughput and memory capacity brings significant opportunities to a wide range of workloads. However, efficiently utilizing these resources remains challen…

cs.DC2026

High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors

Ruimin Shi, Gabin Schieffer, Pei-Hung Lin +3

ARM SVE and RISC-V RVV are emerging vector architectures in high-end processors that support vectorization of flexible vector length. In this work, we leverage an important workloa…

cs.DC2025

Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs

Jacob Wahlgren, Gabin Schieffer, Ruimin Shi +4

Discrete GPUs are a cornerstone of HPC and data center systems, requiring management of separate CPU and GPU memory spaces. Unified Virtual Memory (UVM) has been proposed to ease t…

cs.DC2025

Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive

Gabin Schieffer, Jacob Wahlgren, Ruimin Shi +4

The ever-increasing compute performance of GPU accelerators drives up the need for efficient data movements within HPC applications to sustain performance. Proposed as a solution t…

cs.DC2025

ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace

Ruimin Shi, Gabin Schieffer, Maya Gokhale +3

Vector architectures are essential for boosting computing throughput. ARM provides SVE as the next-generation length-agnostic vector extension beyond traditional fixed-length SIMD.…

quant-ph2025

Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits

Gabin Schieffer, Stefano Markidis, Ivy Peng

Quantum computer simulators are an indispensable tool for prototyping quantum algorithms and verifying the functioning of existing quantum computer hardware. The current largest qu…