3 papers
cs.LG2025
The Anatomy of a Triton Attention Kernel
Burkhard Ringlein, Jan van Lunteren, Radu Stoica +1
A long-standing goal in both industry and academia is to develop an LLM inference platform that is portable across hardware architectures, eliminates the need for low-level hand-tu…
cs.AR2025
GPU Performance Portability needs Autotuning
Burkhard Ringlein, Thomas Parnell, Radu Stoica
As LLMs grow in complexity, achieving state-of-the-art performance requires tight co-design across algorithms, software, and hardware. Today's reliance on a single dominant platfor…
cs.DC2025
Efficient and Reuseable Cloud Configuration Search Using Discovery Spaces
Michael Johnston, Burkhard Ringlein, Christoph Hagleitner +4
Finding the optimal set of cloud resources to deploy a given workload at minimal cost while meeting a defined service level agreement is an active area of research. Combining tens…