collaborators

6 papers

cs.ET2026

Not Your Usual FFT: QFTFFT via Classical Quantum-Circuit Simulation

Stefano Markidis, Gilbert Netzer, Luca Pennati +2

We introduce QFTFFT, a family of HPC FFT libraries that compute the discrete Fourier transform by executing a quantum Fourier transform (QFT) circuit on classical quan…

cs.DC2026

Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors

Ruimin Shi, Maya Gokhale, Pei-Hung Lin +2

The RISC-V Vector Extension~(RVV) is a cornerstone for supporting compute throughout in scientific and machine learning workloads. Yet compiler support and performance monitoring o…

cs.DC2026

Communication Offloading on SmartNIC DPUs: A Quantitative Approach

Jacob Wahlgren, Andong Hu, Roger Pearce +2

SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface. In this wor…

cs.DC2026

Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading

Gabin Schieffer, Ruimin Shi, Jie Ren +1

Advances in GPU compute throughput and memory capacity brings significant opportunities to a wide range of workloads. However, efficiently utilizing these resources remains challen…

cs.DC2026

Making Room for AI: Multi-GPU Molecular Dynamics with Deep Potentials in GROMACS

Luca Pennati, Andong Hu, Ivy Peng +2

GROMACS is a de-facto standard for classical Molecular Dynamics (MD). The rise of AI-driven interatomic potentials that pursue near-quantum accuracy at MD throughput now poses a si…

cs.DC2026

High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors

Ruimin Shi, Gabin Schieffer, Pei-Hung Lin +3

ARM SVE and RISC-V RVV are emerging vector architectures in high-end processors that support vectorization of flexible vector length. In this work, we leverage an important workloa…