3 papers
cs.LG2026
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to mod…
cs.AI2026
Imperfect World Models are Exploitable
Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel +2
We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred o…
cs.AI2025
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill
Human-designed reward functions for reinforcement learning (RL) agents are frequently misaligned with the humans' true, unobservable objectives, and thus act only as proxies. Optim…