2 papers
cs.DC2026
RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM Inference
Guotao Yang, Mingxi Zhao, Haopeng Li +4
Large language model (LLM) agents repeatedly process long, privacy-sensitive contexts, while cloud-only deployment exposes user data beyond the trusted endpoint and fully local dep…
cs.LG2025
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
Zhengchao Wang, Yitao Hu, Jianing Ye +4
Retrieval-Augmented Generation (RAG) is a critical paradigm for building reliable, knowledge-intensive Large Language Model (LLM) applications. However, the multi-stage pipeline (r…