1 paper
Chenyang Ai, Yixing Zhang, Haoran Wu +3
Large language model (LLM) decoding is a major inference bottleneck because its low arithmetic intensity makes performance highly sensitive to memory bandwidth. 3D-stacked near-mem…