14 papers · 1 filter
Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms
Lannie Dalton Hough, Emir Gencer, Hoffmann Muki +1
High-performance collective communication primitives are necessary for a variety of high performance computing (HPC) and machine learning (ML) workloads. State-of-the-art collectiv…
Understanding and Improving Communication Performance in Multi-node LLM Inference
Prajwal Singhania, Siddharth Singh, Lannie Dalton Hough +4
As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only acr…
The Big Send-off: Scalable and Performant Collectives for Deep Learning
Siddharth Singh, Keshav Pradeep, Mahua Singh +2
Collective communication is becoming increasingly important in data center and supercomputer workloads with an increase in distributed AI related jobs. However, existing libraries…
Characterizing Production GPU Workloads using System-wide Telemetry Data
Onur Cankur, Brian Austin, Dhruva Kulkarni +1
GPGPU-accelerated clusters and supercomputers are central to modern high-performance computing (HPC). Over the past decade, these systems continue to expand, and GPUs now expose a…
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
Daniel Nichols, Konstantinos Parasyris, Charles Jekel +2
Language models are now prevalent in software engineering with many developers using them to automate tasks and accelerate their development. While language models have been tremen…
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
Joshua H. Davis, Daniel Nichols, Ishan Khillan +1
GPGPU architectures have become significantly more diverse in recent years, which has led to an emergence of a variety of specialized programming models and software stacks to supp…