3 papers
cs.AI2026
The Impossibility of Eliciting Latent Knowledge
Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport +2
Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users. Consequently, a desirable property…
cs.AI2025
Incentives for Responsiveness, Instrumental Control and Impact
Ryan Carey, Eric Langlois, Chris van Merwijk +2
We introduce three concepts that describe an agent's incentives: response incentives indicate which variables in the environment, such as sensitive demographic information, affect…
cs.AI2025
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…