Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
Qiuyang Zhang, Kai Zhou, Ding Tang +5
Large language models encounter critical GPU memory capacity constraints during long-context inference, where KV cache memory consumption severely limits decode batch sizes. While…
cs.LG2024
Hammer: Towards Efficient Hot-Cold Data Identification via Online Learning
Kai Lu, Siqi Zhao, Jiguang Wan
Efficient management of storage resources in big data and cloud computing environments requires accurate identification of data's "cold" and "hot" states. Traditional methods, such…