6 citations · 6 across the 7 of their papers we have counts for
13 papers · 1 filter
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
Benedikt Hornig, Reuth Mirsky
In shared autonomy, a critical tension arises when an automated assistant must choose between obeying a human's instruction and deliberately overriding it to prevent harm. This saf…
Proceedings of the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind
Nitay Alon, Joseph M. Barnby, Reuth Mirsky +1
This volume includes a selection of papers presented at the 2nd Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2026 in Singapore on 26th January…
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…
GRAIL: Goal Recognition Alignment through Imitation Learning
Osher Elhadad, Felipe Meneguzzi, Reuth Mirsky
Understanding an agent's goals from its behavior is fundamental to aligning AI systems with human intentions. Existing goal recognition methods typically rely on an optimal goal-or…
General Dynamic Goal Recognition using Goal-Conditioned and Meta Reinforcement Learning
Osher Elhadad, Owen Morrissey, Reuth Mirsky
Understanding an agent's goal through its behavior is a common AI problem called Goal Recognition (GR). This task becomes particularly challenging in dynamic environments where goa…
Artificial Intelligent Disobedience: Rethinking the Agency of Our Artificial Teammates
Reuth Mirsky
Artificial intelligence has made remarkable strides in recent years, achieving superhuman performance across a wide range of tasks. Yet despite these advances, most cooperative AI…