agentic learning 1contextual bandits 1large language models 1memory management 1reinforcement learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
The paper introduces MemCon, a framework that treats memory operations of large language model agents as a controllable Markov Decision Process, learning adaptive policies for when…
cs.CL2026
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
Dong Liu, Yanxuan Yu, Ying Nian Wu
The success of large language models (LLMs) across diverse NLP tasks has elevated the importance of reasoning chain optimization as a critical step in aligning model behavior with…