Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…
cs.AI2026
AnyEdit++: Adaptive Long-Form Knowledge Editing via Bayesian Surprise
Bowen Tian, Caixue He, Jiemin Wu +4
Editing complex, long-form knowledge in Large Language Models remains a significant challenge due to the difficulty of maintaining generation coherence. Existing autoregressive met…