1 paper
Yihan Yin, Yinlun Zhao, Zhixin Yun +8
LLM serving is increasingly constrained by memory capacity as model weights, KV caches, and the number of served model variants continue to grow. This report examines High-Bandwidt…