4 papers
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
Ahmed J. Abdelmaksoud, Cristian Sestito, Shiwei Wang +1
The performance gains obtained by large language models (LLMs) are closely linked to their substantial computational and memory requirements. Quantized LLMs offer significant advan…
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
Ahmed J. Abdelmaksoud, Cristian Sestito, Shiwei Wang +1
Transformers are at the core of modern AI nowadays. They rely heavily on matrix multiplication and require efficient acceleration due to their substantial memory and computational…
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
Ahmed J. Abdelmaksoud, Shady Agwa, Themis Prodromakis
Transformers are gaining increasing attention across Natural Language Processing (NLP) application domains due to their outstanding accuracy. However, these data-intensive models a…
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
Cristian Sestito, Ahmed J. Abdelmaksoud, Shady Agwa +1
The Von Neumann bottleneck, which relates to the energy cost of moving data from memory to on-chip core and vice versa, is a serious challenge in state-of-the-art AI architectures,…