2 papers
cs.CL2025
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
Erwei Wang, Samuel Bayliss, Andra Bisca +19
General-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures…
cs.SE2025
Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface
Erika Hunhoff, Joseph Melber, Kristof Denolf +8
Accelerators such as neural processing units (NPUs) deliver an enticing balance of performance and efficiency compared to general purpose compute architectures. However, effectivel…