6 papers
Splyce: SIMD Vectorization of Sparse Coiteration
Kabilan Mahathevan, Poorna Gunathilaka, Kirshanthan Sundararajah
Sparse tensor contractions are bottlenecked by sparse-sparse coiteration loops that resist standard loop vectorization. We present Splyce, an auto-vectorization framework in MLIR t…
A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU
Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah +1
CPU-GPU coscheduling enables simultaneous execution of an application across both processing units, but its efficiency depends on workload partitioning and memory architecture. Thi…
nomp: A Framework for Building Domain Specific Compilers
Thilina Ratnayaka, Kaushik Kulkarni, Nipuna Fernando +8
The low-level GPU programming models (CUDA, HIP, OpenCL, etc.) provide detailed control of the data flow and execution plan of a program in order to extract close-to-metal performa…
Assessing Large Language Models for Stabilizing Numerical Expressions in Scientific Software
Tien Nguyen, Kirshanthan Sundararajah, Muhammad Ali Gulzar
Scientific software relies on high-precision computation, yet finite floating-point representations introduce precision errors that propagate in safety-critical domains. Despite gr…
Eliminate Branches by Melding IR Instructions
Yuze Li, Srinivasan Ramachandra Sharma, Charitha Saumya +2
Branch mispredictions cause catastrophic performance penalties in modern processors, leading to performance loss. While hardware predictors and profile-guided techniques exist, dat…
Bring Your Own Formats and Kernels: Composable Abstractions for Sparse Matrix Computation
Pratyush Das, Amirhossein Basareh, Artem Pelenitsyn +3
Real-world sparse matrices often feature multiple forms of structured sparsity -- rectangular dense blocks, diagonal bands, and scattered entries -- that no single storage format c…