5 papers
SISA: A Scale-In Systolic Array for GEMM Acceleration
Luigi Altamura, Alessio Cicero, Mateo Vázquez Maceiras +2
The currently dominant AI/ML workloads, such as Large Language Models (LLMs), rely on the efficient execution of General Matrix-Matrix Multiplication (GEMM) operations. Thus, most…
Low Latency GNN Accelerator for Quantum Error Correction
Alessio Cicero, Luigi Altamura, Moritz Lange +2
Quantum computers can solve selected problems more efficiently than classical computers, but current devices are limited by high physical error rates. Quantum Error Correction (QEC…
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso
Convolutional Neural Networks (CNNs) serve various applications with diverse performance and resource requirements. Model-aware CNN accelerators best address these diverse requirem…
Simulation of Quantum Computers: Review and Acceleration Opportunities
Alessio Cicero, Mohammad Ali Maleki, Muhammad Waqar Azhar +2
Quantum computing has the potential to revolutionize multiple fields by solving complex problems that can not be solved in reasonable time with current classical computers. Neverth…
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
Fareed Qararyah, Muhammad Waqar Azhar, Mohammad Ali Maleki +1
Depthwise and pointwise convolutions have fewer parameters and perform fewer operations than standard convolutions. As a result, they have become increasingly used in various compa…