3 papers
cs.AR2025
Learnable Sparsification of Die-to-Die Communication via Spike-Based Encoding
Joshua Nardone, Ruijie Zhu, Joseph Callenes +3
Efficient communication is central to both biological and artificial intelligence (AI) systems. In biological brains, the challenge of long-range communication across regions is ad…
cs.AR2025
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
Jinendra Malekar, Peyton Chandarana, Md Hasibul Amin +2
In this paper, we propose PIM-LLM, a hybrid architecture developed to accelerate 1-bit large language models (LLMs). PIM-LLM leverages analog processing-in-memory (PIM) architectur…
cs.AR2025
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
Deepak Vungarala, Mohammed E. Elbtity, Sumiya Syed +5
The increasing complexity and scale of Deep Neural Networks (DNNs) necessitate specialized tensor accelerators, such as Tensor Processing Units (TPUs), to meet various computationa…