1 paper
Xinyu Wang, Yalong Xue, Xiaotian Sun +5
Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flas…