Publications (9)
Open Problems in Frontier AI Risk Management
Marta Ziosi, Miro Plueckebaum, Stephen Casper +26
Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consensus resulting from the rapid pa…
Counterfactual Planning in AGI Systems
Koen Holtman
We present counterfactual planning as a design approach for creating a range of safety mechanisms that can be applied in hypothetical future AI systems which have Artificial Genera…
AGI Agent Safety by Iteratively Improving the Utility Function
Koen Holtman
While it is still unclear if agents with Artificial General Intelligence (AGI) could ever be built, we can already use mathematical models to investigate potential safety systems f…
Demanding and Designing Aligned Cognitive Architectures
Koen Holtman
With AI systems becoming more powerful and pervasive, there is increasing debate about keeping their actions aligned with the broader goals and needs of humanity. This multi-discip…
Corrigibility with Utility Preservation
Koen Holtman
Corrigibility is a safety property for artificially intelligent agents. A corrigible agent will not resist attempts by authorized parties to alter the goals and constraints that we…
Appcessory Economics: Enabling loosely coupled hardware / software innovation
Koen Holtman
An appcessory (app + accessory) is a smart phone accessory that is combined with a specially written app to perform a useful function. An example is a toy helicopter controlled by…