3 papers
cs.AR2024
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
Kosmas Alexandridis, Christodoulos Peltekis, Dionysios Filippas +1
The widespread adoption of machine learning algorithms necessitates hardware acceleration to ensure efficient performance. This acceleration relies on custom matrix engines that op…
cs.AR2024
Error Checking for Sparse Systolic Tensor Arrays
Christodoulos Peltekis, Dionysios Filippas, Giorgos Dimitrakopoulos
Structured sparsity is an efficient way to prune the complexity of modern Machine Learning (ML) applications and to simplify the handling of sparse data in hardware. In such cases,…
cs.AR2023
The Case for Asymmetric Systolic Array Floorplanning
C. Peltekis, D. Filippas, G. Dimitrakopoulos +1
The widespread proliferation of deep learning applications has triggered the need to accelerate them directly in hardware. General Matrix Multiplication (GEMM) kernels are elementa…