activity
20232025
most citedMaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.ARShow all

5 papers · 1 filter

cs.AR2025

Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs

Endri Taka, Andre Roesti, Joseph Melber +3

The high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen…

cs.AR2025

GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines

Kaustubh Mhatre, Endri Taka, Aman Arora

General matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM wo…

cs.AR2025

Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration

Endri Taka, Ning-Chi Huang, Chi-Chih Chang +3

FPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic rese…

cs.AR2024

Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs

Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer +2

FPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading…

cs.AR20231 cited

MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine

Endri Taka, Aman Arora, Kai-Chiang Wu +1

The increasing computational and memory requirements of Deep Learning (DL) workloads has led to outstanding innovations in hardware architectures. An archetype of such architecture…