2 papers
cs.CL2025
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
Dinghong Song, Yuan Feng, Yiwei Wang +6
Large Language Models (LLMs) are widely used in generative applications such as chatting, code generation, and reasoning. However, many realworld workloads such as classification,…
cs.PF2024
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
Shangye Chen, Jin Huang, Shuangyan Yang +7
Tiered memory, built upon a combination of fast memory and slow memory, provides a cost-effective solution to meet ever-increasing requirements from emerging applications for large…