11 papers
LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
Yimin Wang, Yue Jiet Chong, Xuanyao Fong
LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promisi…
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration
Yue Jiet Chong, Yimin Wang, Zhen Wu +3
Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high ut…
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Yue Jiet Chong, Yimin Wang, Wei Zhang +1
LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge dev…
Continuous Optimization for Satisfiability Modulo Theories on Linear Real Arithmetic
Yunuo Cen, Daniel Ebler, Xuanyao Fong
Efficient solutions for satisfiability modulo theories (SMT) are integral in industrial applications such as hardware verification and design automation. Existing approaches are pr…
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
Abhishek Tyagi, Yunuo Cen, Shrey Dhorajiya +2
Large Language Models (LLMs) exhibit substantial parameter redundancy, particularly in Feed-Forward Networks (FFNs). Existing pruning methods suffer from two primary limitations. F…
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
Yue Jiet Chong, Yimin Wang, Zhen Wu +1
This paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM…