2 papers
cs.CL2026
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
Rongsheng Zhang, Ruofan Hu, Weijie Chen +7
While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Current systems typically rely…
cs.LG2026
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
Zhaochun Li, Chen Wang, Jionghao Bai +4
The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…