1 paper
Hongshi Tan, Yao Chen, Gustavo Alonso +2
Large language models (LLMs) now scale to trillions of parameters, driving weight storage into the terabyte regime and creating an acute mismatch with GPU memory capacity. Although…