18 papers
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
Ram Rachum, Yotam Amitai, Bálint Gyevnár +2
This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…
HumAIN: Human-Aware Implicit Social Robot Navigation
Daeun Song, Nhat Le, Jeffrey Chen +7
Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orientation. We present Human-Aware Implicit Soc…
BXRL: Behavior-Explainable Reinforcement Learning
Ram Rachum, Yotam Amitai, Yonatan Nakar +2
A major challenge of Reinforcement Learning is that agents often learn undesired behaviors that seem to defy the reward structure they were given. Explainable Reinforcement Learnin…
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
Benedikt Hornig, Reuth Mirsky
In shared autonomy, a critical tension arises when an automated assistant must choose between obeying a human's instruction and deliberately overriding it to prevent harm. This saf…
Proceedings of the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind
Nitay Alon, Joseph M. Barnby, Reuth Mirsky +1
This volume includes a selection of papers presented at the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2026 in Singapore on 26th January…
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…