2 papers
cs.DC2026
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
Wenfeng Wang, Xiaofeng Hou, Peng Tang +5
Retrieval-Augmented Generation (RAG) systems enhance the performance of large language models (LLMs) by incorporating supplementary retrieved documents, enabling more accurate and…
cs.DC2025
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
Jing Wang, Chao Li, Taolei Wang +4
The growing scale of data requires efficient memory subsystems with large memory capacity and high memory performance. Disaggregated architecture has become a promising solution fo…