1 citations · 1 across the 3 of their papers we have counts for
5 papers
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
Ahmed J. Abdelmaksoud, Cristian Sestito, Shiwei Wang +1
The performance gains obtained by large language models (LLMs) are closely linked to their substantial computational and memory requirements. Quantized LLMs offer significant advan…
A flexible language model-assisted electronic design automation framework
Cristian Sestito, Panagiota Kontou, Pratibha Verma +5
Large language models (LLMs) are transforming electronic design automation (EDA) by enhancing design stages such as schematic design, simulation, netlist synthesis, and place-and-r…
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
Ahmed J. Abdelmaksoud, Cristian Sestito, Shiwei Wang +1
Transformers are at the core of modern AI nowadays. They rely heavily on matrix multiplication and require efficient acceleration due to their substantial memory and computational…
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
Cristian Sestito, Ahmed J. Abdelmaksoud, Shady Agwa +1
The Von Neumann bottleneck, which relates to the energy cost of moving data from memory to on-chip core and vice versa, is a serious challenge in state-of-the-art AI architectures,…
L-Sort: On-chip Spike Sorting with Efficient Median-of-Median Detection and Localization-based Clustering
Yuntao Han, Yihan Pan, Xiongfei Jiang +4
Spike sorting is a critical process for decoding large-scale neural activity from extracellular recordings. The advancement of neural probes facilitates the recording of a high num…