3 papers
cs.CL2026
TokenPilot: Cache-Efficient Context Management for LLM Agents
Buqiang Xu, Zirui Xue, Dianmou Chen +12
As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize…
cs.CL2025
Tuning LLMs by RAG Principles: Towards LLM-native Memory
Jiale Wei, Shuchi Wu, Ruochen Liu +3
Memory, additional information beyond the training of large language models (LLMs), is crucial to various real-world applications, such as personal assistant. The two mainstream so…
cs.AI2025
AI-native Memory 2.0: Second Me
Jiale Wei, Xiang Ying, Tao Gao +3
Human interaction with the external world fundamentally involves the exchange of personal memory, whether with other individuals, websites, applications, or, in the future, AI agen…