3 papers
cs.AR2026
Accelerating CRONet on AMD Versal AIE-ML Engines
Kaustubh Mhatre, Vedant Tewari, Aditya Ray +4
Topology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while sa…
cs.AR2025
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
Kaustubh Mhatre, Endri Taka, Aman Arora
General matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM wo…
cs.AR2023
PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware Computation
Aman Arora, Jian Weng, Siyuan Ma +2
Bit-serial Processing-In-Memory (PIM) is an attractive paradigm for accelerator architectures, for parallel workloads such as Deep Learning (DL), because of its capability to achie…