2 papers
cs.AI2026
Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Benjamin Poole, Minwoo Lee
Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feed…
cs.LG2026
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
Benjamin Poole, Andrew Quinn, Li Yang +1
Data rehearsal has emerged as a leading approach for mitigating catastrophic forgetting in Continual Reinforcement Learning (CRL). However, existing work remains confined to policy…