collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2025

A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces

Ankur Lahiry, Ayush Pokharel, Banooqa Banday +5

Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume an…

cs.DC2025

PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training

Seth Ockerman, Amal Gueroudji, Tanwi Mallick +4

Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to sm…

cs.DC2025

Automatic Metadata Capture and Processing for High-Performance Workflows

Polina Shpilker, Line Pouchard

Modern workflows run on increasingly heterogeneous computing architectures and with this heterogeneity comes additional complexity. We aim to apply the FAIR principles for research…

cs.DC2025

Scalable GPU Performance Variability Analysis framework

Ankur Lahiry, Ayush Pokharel, Seth Ockerman +3

Analyzing large-scale performance logs from GPU profilers often requires terabytes of memory and hours of runtime, even for basic summaries. These constraints prevent timely insigh…

cs.DC2024

Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows

Ozgur Ozan Kilic, Tianle Wang, Matteo Turilli +4

Workflows are critical for scientific discovery. However, the sophistication, heterogeneity, and scale of workflows make building, testing, and optimizing them increasingly challen…