9 papers · 1 filter
Understanding and Improving Communication Performance in Multi-node LLM Inference
Prajwal Singhania, Siddharth Singh, Lannie Dalton Hough +4
As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only acr…
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
Daniel Nichols, Konstantinos Parasyris, Caetano Melone +3
As high-performance computing and AI workloads become increasingly dependent on GPUs, maintaining high performance across rapidly evolving hardware generations has become a major c…
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
Gregory Bolet, Giorgis Georgakoudis, Konstantinos Parasyris +4
Modern GPU software stacks demand developers who can anticipate performance bottlenecks before ever launching a kernel; misjudging floating-point workloads upstream can derail tuni…
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
Daniel Nichols, Konstantinos Parasyris, Charles Jekel +2
Language models are now prevalent in software engineering with many developers using them to automate tasks and accelerate their development. While language models have been tremen…
Taking GPU Programming Models to Task for Performance Portability
Joshua H. Davis, Pranav Sivaraman, Joy Kitson +5
Portability is critical to ensuring high productivity in developing and maintaining scientific software as the diversity in on-node hardware architectures increases. While several…
Can Large Language Models Predict Parallel Code Performance?
Gregory Bolet, Giorgis Georgakoudis, Harshitha Menon +5
Accurate determination of the performance of parallel GPU code typically requires execution-time profiling on target hardware -- an increasingly prohibitive step due to limited acc…