activity
20242026
collaborators

18 papers

cs.LG2026

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

Ram Rachum, Yotam Amitai, Bálint Gyevnár +2

This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…

cs.RO2026

HumAIN: Human-Aware Implicit Social Robot Navigation

Daeun Song, Nhat Le, Jeffrey Chen +7

Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orientation. We present Human-Aware Implicit Soc…

cs.LG2026

BXRL: Behavior-Explainable Reinforcement Learning

Ram Rachum, Yotam Amitai, Yonatan Nakar +2

A major challenge of Reinforcement Learning is that agents often learn undesired behaviors that seem to defy the reward structure they were given. Explainable Reinforcement Learnin…

cs.AI2026

The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes

Benedikt Hornig, Reuth Mirsky

In shared autonomy, a critical tension arises when an automated assistant must choose between obeying a human's instruction and deliberately overriding it to prevent harm. This saf…

cs.AI2026

Proceedings of the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind

Nitay Alon, Joseph M. Barnby, Reuth Mirsky +1

This volume includes a selection of papers presented at the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2026 in Singapore on 26th January…

cs.AI2026

Agents of Chaos

Natalie Shapira, Chris Wendler, Avery Yen +35

We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…