Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
Guibin Zhang, Xun Xu, Yanwei Yue +4
Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning how to evolve through it. Existing memory agents can store tra…
cs.CL2025
Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny
Chuanhao Yan, Fengdi Che, Xuhan Huang +12
Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant challenge: their verification proce…