5 papers
AI Integrity: Defending Against Backdoors and Secret Loyalties
Dave Banerjee, Onni Aarne
AI integrity means ensuring AI systems are free from secret or unauthorized modifications that could compromise their behavior. Integrity represents one pillar of the confidentiali…
International Security Applications of Flexible Hardware-Enabled Guarantees
Onni Aarne, James Petrie
As AI capabilities advance rapidly, flexible hardware-enabled guarantees (flexHEGs) offer opportunities to address international security challenges through comprehensive governanc…
Flexible Hardware-Enabled Guarantees for AI Compute
James Petrie, Onni Aarne, Nora Ammann +1
As artificial intelligence systems become increasingly powerful, they pose growing risks to international security, creating urgent coordination challenges that current governance…
Technical Options for Flexible Hardware-Enabled Guarantees
James Petrie, Onni Aarne
Frontier AI models pose increasing risks to public safety and international security, creating a pressing need for AI developers to provide credible guarantees about their developm…
Open Problems in Technical AI Governance
Anka Reuel, Ben Bucknall, Stephen Casper +30
AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated. In many cases, the barriers and uncertainties faced are at…