3 papers
cs.LG2026
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
Zizhe Chen, Jiqian Dong, Yizhou Tian +4
Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state value estimation is critical for…
cs.CL2026
RecMem: Recurrence-based Memory Consolidation for Efficient and Effective Long-Running LLM Agents
Zijie Dai, Shiyuan Deng, Sheng Guan +4
Memory systems often organize user-agent interactions as retrievable external memory and are crucial for long-running agents by overcoming the limited context windows of LLMs. Howe…
cs.LG2026
Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
Zixuan Chen, Hao Lin, Zizhe Chen +6
LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply rather than correct. We term this…