Showing cs.ARShow all
3 papers · 1 filter
cs.AR2025
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
João Paulo Cardoso de Lima, Marc Dietrich, Jeronimo Castrillon +1
Structured sparsity enables deploying large language models (LLMs) on resource-constrained systems. Approaches like dense-to-sparse fine-tuning are particularly compelling, achievi…
cs.AR2025
Count2Multiply: Reliable In-Memory High-Radix Counting
João Paulo Cardoso de Lima, Benjamin Franklin Morris, Asif Ali Khan +2
Computing-in-memory (CIM) has been demonstrated across various memory technologies, ranging from memristive crossbars performing analog dot-product computations to large-scale digi…
cs.AR2025
Modeling and Simulating Emerging Memory Technologies: A Tutorial
Yun-Chih Chen, Tristan Seidl, Nils Hölscher +15
Non-volatile Memory (NVM) technologies present a promising alternative to traditional volatile memories such as SRAM and DRAM. Due to the limited availability of real NVM devices,…