1 paper
Wanning Zhang, Tongzhou Gu, Marco Canini +2
Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level c…