2 papers
cs.AR2026
Cache-Resident LLM Inference in GB-Scale Last-Level Caches
Wanning Zhang, Tongzhou Gu, Marco Canini +2
Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level c…
cs.LG2026
FOCUS: DLLMs Know How to Tame Their Compute Bound
Kaihua Liang, Xin Tan, An Zhong +2
Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost. In this work, we identi…