1 paper
Xinjun Yang, Qingda Hu, Junru Li +9
The rapid increase in LLM model sizes and the growing demand for long-context inference have made memory a critical bottleneck in GPU-accelerated serving systems. Although high-ban…