3 papers
cs.SE2026
Kerncap: Automated Kernel Extraction and Isolation for AMD GPUs
Cole Ramos, Keith Lowery
Iterative GPU kernel tuning is bottlenecked by the scale of the applications that host the kernels. Rapid iteration requires isolating the kernel so it can be edited, recompiled, a…
cs.DC2025
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
Arya Tschand, Muhammad Awad, Ryan Swann +5
Large language models (LLMs) have shown progress in GPU kernel performance engineering using inefficient search-based methods that optimize around runtime. Any existing approach la…
cs.LG2025
Omniwise: Predicting GPU Kernels Performance with LLMs
Zixian Wang, Cole Ramos, Muhammad A. Awad +1
In recent years, the rapid advancement of deep neural networks (DNNs) has revolutionized artificial intelligence, enabling models with unprecedented capabilities in understanding,…