collaborators

9 papers

cs.LG2026

Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

Victor Akinwande, J. Zico Kolter, Aran Nayebi

Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidenc…

cs.RO2026

Tactile Genesis: Exploring Tactile Sensors at Scale for Learning Dexterous Tasks

Trinity Chung, Kashu Yamazaki, Dhruv Patel +4

Tactile sensing is critical for contact-rich dexterous manipulation, yet it remains unclear which tactile abstractions a policy needs and when richer tactile fields justify their h…

cs.LG2026

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

Aran Nayebi

As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Classical results show that optimal contro…

cs.LG2026

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

Jeremy Tien, Abishek Anand, Yu-Rou Tuan +3

As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding t…

econ.GN2025

When Do AI Gains Become Broadly Shareable? A Policy Threshold for AI-Driven Automation

Aran Nayebi

AI-driven automation generates broad-based social benefit only if technical gains become visible, durable, and publicly claimable. We develop a policy-facing stress test by extendi…

cs.AI2025

Core Safety Values for Provably Corrigible Agents

Aran Nayebi

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework cons…