5 citations · 13 across the 14 of their papers we have counts for
16 papers · 1 filter
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
Elena Baskakova, William Bergeron, Matthew Hubbell +2
Supercomputers are complex, dynamic systems that serve thousands of users and are built with thousands of compute nodes. Due to the vast amounts of system and performance data need…
Easy Acceleration with Distributed Arrays
Jeremy Kepner, Chansup Byun, LaToya Anderson +20
High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (acros…
GPU Sharing with Triples Mode
Chansup Byun, Albert Reuther, LaToya Anderson +19
There is a tremendous amount of interest in AI/ML technologies due to the proliferation of generative AI applications such as ChatGPT. This trend has significantly increased demand…
Supercomputer 3D Digital Twin for User Focused Real-Time Monitoring
William Bergeron, Matthew Hubbell, Daniel Mojica +17
Real-time supercomputing performance analysis is a critical aspect of evaluating and optimizing computational systems in a dynamic user environment. The operation of supercomputers…
HPC with Enhanced User Separation
Andrew Prout, Albert Reuther, Michael Houle +19
HPC systems used for research run a wide variety of software and workflows. This software is often written or modified by users to meet the needs of their research projects, and ra…
LLload: Simplifying Real-Time Job Monitoring for HPC Users
Chansup Byun, Julia Mullen, Albert Reuther +16
One of the more complex tasks for researchers using HPC systems is performance monitoring and tuning of their applications. Developing a practice of continuous performance improvem…