2 papers
cs.LG2025
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
Krishna Teja Chitty-Venkata, Jie Ye, Xian-He Sun +4
KV caching significantly improves the efficiency of Large Language Model (LLM) inference by storing attention states from previously processed tokens, enabling faster generation of…
cs.DC2025
Performance Models for a Two-tiered Storage System
Aparna Sasidharan, Xian-He, Jay Lofstead +1
This work describes the design, implementation and performance analysis of a distributed two-tiered storage software. The first tier functions as a distributed software cache imple…