most citedMaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.ARShow all

5 papers · 1 filter

cs.AR2025

CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs

Jiajun Hu, Chetan Choppali Sudarshan, Vidya A. Chhabria +1

Over the years, the chip industry has consistently developed high-performance processors to address the increasing demands across diverse applications. However, the rapid expansion…

cs.AR2024

Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs

Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer +2

FPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading…

cs.AR2023

PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware Computation

Aman Arora, Jian Weng, Siyuan Ma +2

Bit-serial Processing-In-Memory (PIM) is an attractive paradigm for accelerator architectures, for parallel workloads such as Deep Learning (DL), because of its capability to achie…

cs.AR20231 cited

MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine

Endri Taka, Aman Arora, Kai-Chiang Wu +1

The increasing computational and memory requirements of Deep Learning (DL) workloads has led to outstanding innovations in hardware architectures. An archetype of such architecture…

cs.AR2023

ULEEN: A Novel Architecture for Ultra Low-Energy Edge Neural Networks

Zachary Susskind, Aman Arora, Igor D. S. Miranda +9

The deployment of AI models on low-power, real-time edge devices requires accelerators for which energy, latency, and area are all first-order concerns. There are many approaches t…