1 paper
Shouxu Lin, Zhiyuan Guo, Jiaxin Lin
LLM inference is constrained by GPU memory capacity and bandwidth. Tiered memory architectures mitigate this by allowing the GPU to offload memory to the remote tier. However, exis…