4 papers
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
João Paulo Cardoso de Lima, Marc Dietrich, Jeronimo Castrillon +1
Structured sparsity enables deploying large language models (LLMs) on resource-constrained systems. Approaches like dense-to-sparse fine-tuning are particularly compelling, achievi…
CoMoNM: A Cost Modeling Framework for Compute-Near-Memory Systems
Hamid Farzaneh, Asif Ali Khan, Jeronimo Castrillon
Compute-Near-Memory (CNM) systems offer a promising approach to mitigate the von Neumann bottleneck by bringing computational units closer to data. However, optimizing for these ar…
Count2Multiply: Reliable In-Memory High-Radix Counting
João Paulo Cardoso de Lima, Benjamin Franklin Morris, Asif Ali Khan +2
Computing-in-memory (CIM) has been demonstrated across various memory technologies, ranging from memristive crossbars performing analog dot-product computations to large-scale digi…
All-in-Memory Stochastic Computing using ReRAM
João Paulo C. de Lima, Mehran Shoushtari Moghadam, Sercan Aygun +3
As the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for handling complex tasks. Stochasti…