3 papers
cs.MA2026
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Chelsea Zou, Yiheng Yao, Selena She +2
Personal AI assistants are beginning to act as delegates with access to calendars, inboxes, and user preferences. Calendar scheduling makes the trust problem concrete: an assistant…
cs.AI2026
Measuring Progress Toward AGI: A Cognitive Framework
Ryan Burnell, Yumeya Yamamori, Orhan Firat +10
Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, an…
cs.AI2025
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…