activity
20242026
collaborators

7 papers

cs.LG2026

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

Ram Rachum, Yotam Amitai, Bálint Gyevnár +2

This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…

cs.LG2026

BXRL: Behavior-Explainable Reinforcement Learning

Ram Rachum, Yotam Amitai, Yonatan Nakar +2

A major challenge of Reinforcement Learning is that agents often learn undesired behaviors that seem to defy the reward structure they were given. Explainable Reinforcement Learnin…

cs.LG2026

Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction

Tianyi Alex Qiu, Micah Carroll, Cameron Allen

The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating fr…

cs.LG2025

Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects

Jonathan Colaço Carr, Qinyi Sun, Cameron Allen

Skills are essential for unlocking higher levels of problem solving. A common approach to discovering these skills is to learn ones that reliably reach different states, thus empow…

cs.LG2025

From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning

Rafael Rodriguez-Sanchez, Cameron Allen, George Konidaris

Algorithms that exploit factored Markov decision processes are far more sample-efficient than factor-agnostic methods, yet they assume a factored representation is known a priori -…

cs.LG2025

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains

Ruo Yu Tao, Kaicheng Guo, Cameron Allen +1

Mitigating partial observability is a necessary but challenging task for general reinforcement learning algorithms. To improve an algorithm's ability to mitigate partial observabil…