1 paper
Yuning Han, Yangchenchen Jin, Dylan Zhao +1
Auto-regressive decoding in Large Language Models (LLMs) is inherently memory-bound: every generation step requires loading the model weights and intermediate results from memory (…