activity
20192026
most citedAn analytic performance model for overlapping execution of memory-bound loop kernels on multicore CPUs

2 citations · 2 across the 5 of their papers we have counts for

collaborators

8 papers

cs.DC2026

Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters

Ayesha Afzal, Georg Hager, Gerhard Wellein

The escalating computational demands and energy footprint of GPU-accelerated computing systems complicate informed design and operational decisions. We present the first release of…

cs.DC2025

GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs

Ayesha Afzal, Anna Kahler, Georg Hager +1

Molecular dynamics simulations are essential tools in computational biophysics, but their performance depend heavily on hardware choices and configuration. In this work, we present…

cs.DC2025

Exploring metrics for analyzing dynamic behavior in MPI programs via a coupled-oscillator model

Ayesha Afzal, Georg Hager, Gerhard Wellen

We propose a novel, lightweight, and physically inspired approach to modeling the dynamics of parallel distributed-memory programs. Inspired by the Kuramoto model, we represent MPI…

cs.DC2024

Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters

Ayesha Afzal, Georg Hager, Gerhard Wellein

We present a thorough performance and energy consumption analysis of the LULESH proxy application in its OpenMP and MPI variants on two different clusters based on Intel Ice Lake (…

cs.DC2021

Analytic Modeling of Idle Waves in Parallel Programs: Communication, Cluster Topology, and Noise Impact

Ayesha Afzal, Georg Hager, Gerhard Wellein

Most distributed-memory bulk-synchronous parallel programs in HPC assume that compute resources are available continuously and homogeneously across the allocated set of compute nod…

cs.DC20202 cited

An analytic performance model for overlapping execution of memory-bound loop kernels on multicore CPUs

Ayesha Afzal, Georg Hager, Gerhard Wellein

Complex applications running on multicore processors show a rich performance phenomenology. The growing number of cores per ccNUMA domain complicates performance analysis of memory…