6 papers · 1 filter
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
Sarah Liaw, Benjamin Plaut
In high-stakes AI applications, even a single action can cause irreparable damage. However, nearly all of sequential decision-making theory assumes that all errors are recoverable…
Safety Training May Persist Through Helpfulness Optimization in LLM Agents
Benjamin Plaut
Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step…
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh +1
When deployed in the real world, AI agents will inevitably face challenges that exceed their individual capabilities. A critical component of AI safety is an agent's ability to rec…
Safe Learning Under Irreversible Dynamics via Asking for Help
Benjamin Plaut, Juan Liévano-Karim, Hanlin Zhu +1
Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, w…
Avoiding Catastrophe in Online Learning by Asking for Help
Benjamin Plaut, Hanlin Zhu, Stuart Russell
Most learning algorithms with formal regret guarantees assume that all mistakes are recoverable and essentially rely on trying all possible behaviors. This approach is problematic…
Getting By Goal Misgeneralization With a Little Help From a Mentor
Tu Trinh, Mohamad H. Danesh, Nguyen X. Khanh +1
While reinforcement learning (RL) agents often perform well during training, they can struggle with distribution shift in real-world deployments. One particularly severe risk of di…