2 papers
cs.AR2026
Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems
Yuannuo Feng, Wenyong Zhou, Yuang Ma +5
Analog compute-in-memory (CIM) arrays have emerged as a promising substrate for energy-efficient LLM inference, particularly for weight-stationary computations in linear layers. Ho…
cs.LG2026
SLaB: Sparse-Lowrank-Binary Decomposition for Efficient Large Language Models
Ziwei Li, Yuang Ma, Yi Kang
The rapid growth of large language models (LLMs) presents significant deployment challenges due to their massive computational and memory demands. While model compression, such as…