1 paper
Guoqiang Zou, Wanyu Wang, Hao Zheng +2
Existing memory management techniques severely hinder efficient Large Language Model serving on accelerators constrained by poor random-access bandwidth.While static pre-allocation…