7 papers
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
Ram Rachum, Yotam Amitai, Bálint Gyevnár +2
This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…
BXRL: Behavior-Explainable Reinforcement Learning
Ram Rachum, Yotam Amitai, Yonatan Nakar +2
A major challenge of Reinforcement Learning is that agents often learn undesired behaviors that seem to defy the reward structure they were given. Explainable Reinforcement Learnin…
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
Tianyi Alex Qiu, Micah Carroll, Cameron Allen
The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating fr…
Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects
Jonathan Colaço Carr, Qinyi Sun, Cameron Allen
Skills are essential for unlocking higher levels of problem solving. A common approach to discovering these skills is to learn ones that reliably reach different states, thus empow…
From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning
Rafael Rodriguez-Sanchez, Cameron Allen, George Konidaris
Algorithms that exploit factored Markov decision processes are far more sample-efficient than factor-agnostic methods, yet they assume a factored representation is known a priori -…
Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains
Ruo Yu Tao, Kaicheng Guo, Cameron Allen +1
Mitigating partial observability is a necessary but challenging task for general reinforcement learning algorithms. To improve an algorithm's ability to mitigate partial observabil…