Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
arXiv:2209.09731 · doi:10.1145/3581576.3581621
Abstract
This paper assesses and reports the experience of ten teams working to port,validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems built by GIGABYTE, each one equipped with a server-class Arm CPU from Ampere Computing and A100 data center GPU from NVIDIA Corp. The systems are connected together using Infiniband high-bandwidth low-latency interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.
References in corpus (14)
- Highly Improved Staggered Quarks on the Lattice, with Applications to Charm Physics
- The ARM Scalable Vector Extension
- Accelerating Dynamical Fermion Computations using the Rational Hybrid Monte Carlo (RHMC) Algorithm with Multiple Pseudofermion Fields
- QMCPACK: Advances in the development, efficiency, and application of auxiliary field and real-space variational and diffusion Quantum Monte Carlo
- An assessment of multicomponent flow models and interface capturing schemes for spherical bubble dynamics
- Accelerating Staggered Fermion Dynamics with the Rational Hybrid Monte Carlo (RHMC) Algorithm
- MFC: An open-source high-order multi-component, multi-phase, and multi-scale compressible flow solver
- Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library
- A quantitative comparison of phase-averaged models for bubbly, cavitating flows
- A Smoothed Particle Hydrodynamics Mini-App for Exascale
- Hybrid quadrature moment method for accurate and stable representation of non-Gaussian processes and their dynamics
- Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
- Challenges Porting a C++ Template-Metaprogramming Abstraction Layer to Directive-based Offloading
- First Experiences in Performance Benchmarking with the New SPEChpc 2021 Suites