5 papers
Measuring and mitigating overreliance to build human-compatible AI
Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14
Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…
How are AI agents used? Evidence from 177,000 MCP tools
Merlin Stein
Today's AI agents are built on large language models (LLMs) equipped with tools to access and modify external environments, such as corporate file systems, API-accessible platforms…
Who Should Run Advanced AI Evaluations -- AISIs?
Merlin Stein, Milan Gandhi, Theresa Kriecherbauer +2
Artificial Intelligence (AI) Safety Institutes and governments worldwide are deciding whether they evaluate advanced AI themselves, support a private evaluation ecosystem or do bot…
Measuring AI agent autonomy: Towards a scalable approach with code inspection
Peter Cihon, Merlin Stein, Gagan Bansal +2
AI agents are AI systems that can achieve complex goals autonomously. Assessing the level of agent autonomy is crucial for understanding both their potential benefits and risks. Cu…
Monitoring Human Dependence On AI Systems With Reliance Drills
Rosco Hunter, Richard Moulange, Jamie Bernardi +1
AI systems are assisting humans with increasingly diverse intellectual tasks but are still prone to mistakes. Humans are over-reliant on this assistance if they trust AI-generated…