2 papers
cs.CL2025
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
Huan Yang, Renji Zhang, Mingzhe Huang +5
Recent advances in long-text understanding have pushed the context length of large language models (LLMs) up to one million tokens. It boosts LLMs's accuracy and reasoning capacity…
cs.CR2024
A First Look At Efficient And Secure On-Device LLM Inference Against KV Leakage
Huan Yang, Deyu Zhang, Yudong Zhao +2
Running LLMs on end devices has garnered significant attention recently due to their advantages in privacy preservation. With the advent of lightweight LLM models and specially des…