2 papers
cs.CR2026
A Survey on Split Learning for LLM Fine-Tuning: Models, Systems, and Privacy Optimizations
Zihan Liu, Yizhen Wang, Rui Wang +2
Fine-tuning unlocks large language models (LLMs) for specialized applications, but its high computational cost often puts it out of reach for resource-constrained organizations. Wh…
cs.AR2026
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
Xianzhe Zheng, Zhengheng Wang, Ruiyan Ma +17
The memory-for-computation paradigm of KV caching is essential for accelerating large language model (LLM) inference service, but limited GPU high-bandwidth memory (HBM) capacity m…