3 papers
cs.OS2025
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
Shaoting Feng, Yuhan Liu, Hanchen Li +11
Reusing KV cache is essential for high efficiency of Large Language Model (LLM) inference systems. With more LLM users, the KV cache footprint can easily exceed GPU memory capacity…
cs.AR2025
FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
Ao Shen, Rui Zhang, Junping Zhao
As large language models (LLMs) continue to scale, multi-node deployment has become a necessity. Consequently, communication has become a critical performance bottleneck. Current i…
cs.CL2025
Improving Fairness of Large Language Models in Multi-document Summarization
Haoyuan Li, Rui Zhang, Snigdha Chaturvedi
Fairness in multi-document summarization (MDS) is crucial for providing comprehensive views across documents with diverse social attribute values, which can significantly impact de…