4 papers
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
Ankur Lahiry, Ayush Pokharel, Banooqa Banday +5
Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume an…
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
Seth Ockerman, Amal Gueroudji, Tanwi Mallick +4
Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to sm…
Automatic Metadata Capture and Processing for High-Performance Workflows
Polina Shpilker, Line Pouchard
Modern workflows run on increasingly heterogeneous computing architectures and with this heterogeneity comes additional complexity. We aim to apply the FAIR principles for research…
Scalable GPU Performance Variability Analysis framework
Ankur Lahiry, Ayush Pokharel, Seth Ockerman +3
Analyzing large-scale performance logs from GPU profilers often requires terabytes of memory and hours of runtime, even for basic summaries. These constraints prevent timely insigh…