9 papers
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
Victor Akinwande, J. Zico Kolter, Aran Nayebi
Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidenc…
Tactile Genesis: Exploring Tactile Sensors at Scale for Learning Dexterous Tasks
Trinity Chung, Kashu Yamazaki, Dhruv Patel +4
Tactile sensing is critical for contact-rich dexterous manipulation, yet it remains unclear which tactile abstractions a policy needs and when richer tactile fields justify their h…
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
Aran Nayebi
As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Classical results show that optimal contro…
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
Jeremy Tien, Abishek Anand, Yu-Rou Tuan +3
As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding t…
When Do AI Gains Become Broadly Shareable? A Policy Threshold for AI-Driven Automation
Aran Nayebi
AI-driven automation generates broad-based social benefit only if technical gains become visible, durable, and publicly claimable. We develop a policy-facing stress test by extendi…
Core Safety Values for Provably Corrigible Agents
Aran Nayebi
We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework cons…