4 papers · 1 filter
Learnable Sparsification of Die-to-Die Communication via Spike-Based Encoding
Joshua Nardone, Ruijie Zhu, Joseph Callenes +3
Efficient communication is central to both biological and artificial intelligence (AI) systems. In biological brains, the challenge of long-range communication across regions is ad…
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
Jinendra Malekar, Peyton Chandarana, Md Hasibul Amin +2
In this paper, we propose PIM-LLM, a hybrid architecture developed to accelerate 1-bit large language models (LLMs). PIM-LLM leverages analog processing-in-memory (PIM) architectur…
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
Deepak Vungarala, Mohammed E. Elbtity, Sumiya Syed +5
The increasing complexity and scale of Deep Neural Networks (DNNs) necessitate specialized tensor accelerators, such as Tensor Processing Units (TPUs), to meet various computationa…
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
Mohammed Elbtity, Peyton Chandarana, Ramtin Zand
Tensor processing units (TPUs) are one of the most well-known machine learning (ML) accelerators utilized at large scale in data centers as well as in tiny ML applications. TPUs of…