5 papers
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
Alejandro Ruiz y Mesa, Guilherme Korol, Moritz Riesterer +2
LLM deployment on resource-constrained edge devices faces severe latency constraints, particularly in real-time applications where delayed responses can compromise safety or usabil…
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
João Paulo Cardoso de Lima, Marc Dietrich, Jeronimo Castrillon +1
Structured sparsity enables deploying large language models (LLMs) on resource-constrained systems. Approaches like dense-to-sparse fine-tuning are particularly compelling, achievi…
Count2Multiply: Reliable In-Memory High-Radix Counting
João Paulo Cardoso de Lima, Benjamin Franklin Morris, Asif Ali Khan +2
Computing-in-memory (CIM) has been demonstrated across various memory technologies, ranging from memristive crossbars performing analog dot-product computations to large-scale digi…
All-in-Memory Stochastic Computing using ReRAM
João Paulo C. de Lima, Mehran Shoushtari Moghadam, Sercan Aygun +3
As the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for handling complex tasks. Stochasti…
Modeling and Simulating Emerging Memory Technologies: A Tutorial
Yun-Chih Chen, Tristan Seidl, Nils Hölscher +15
Non-volatile Memory (NVM) technologies present a promising alternative to traditional volatile memories such as SRAM and DRAM. Due to the limited availability of real NVM devices,…