Publications (5)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
Erwei Wang, Samuel Bayliss, Andra Bisca +19
General-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures…
CHARM: Composing Heterogeneous Accelerators for Matrix Multiply on Versal ACAP Architecture
Jinming Zhuang, Jason Lau, Hanchen Ye +10
Dense matrix multiply (MM) serves as one of the most heavily used kernels in deep learning applications. To cope with the high computation demands of these applications, heterogene…
Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface
Erika Hunhoff, Joseph Melber, Kristof Denolf +8
Accelerators such as neural processing units (NPUs) deliver an enticing balance of performance and efficiency compared to general purpose compute architectures. However, effectivel…
SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation
Gagandeep Singh, Alireza Khodamoradi, Kristof Denolf +6
Fast and accurate climate simulations and weather predictions are critical for understanding and preparing for the impact of climate change. Real-world weather and climate modeling…
Comparing Energy Efficiency of CPU, GPU and FPGA Implementations for Vision Kernels
Murad Qasaimeh, Kristof Denolf, Jack Lo +3
Developing high performance embedded vision applications requires balancing run-time performance with energy constraints. Given the mix of hardware accelerators that exist for embe…