4 papers
IRONSmith: A Visual Dataflow Design Environment for AMD Ryzen AI NPUs
Brock Sorenson, Samer Ali, Curt John Bansil +1
Machine learning inference increasingly relies on specialized hardware accelerators for throughput and power efficiency. Neural Processing Units (NPUs), such as the AMD Ryzen AI NP…
Accelerating CRONet on AMD Versal AIE-ML Engines
Kaustubh Mhatre, Vedant Tewari, Aditya Ray +4
Topology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while sa…
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
Kaustubh Mhatre, Endri Taka, Aman Arora
General matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM wo…
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
Endri Taka, Ning-Chi Huang, Chi-Chih Chang +3
FPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic rese…