Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
Qiuyang Zhang, Kai Zhou, Ding Tang +5
Large language models encounter critical GPU memory capacity constraints during long-context inference, where KV cache memory consumption severely limits decode batch sizes. While…
cs.LG2025
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Jiecheng Zhou, Ding Tang, Rong Fu +8
The burgeoning computational demands for training large language models (LLMs) necessitate efficient methods, including quantized training, which leverages low-bit arithmetic opera…