papers

Publications (9)

cs.LG2026

Open Problems in Frontier AI Risk Management

Marta Ziosi, Miro Plueckebaum, Stephen Casper +26

Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consensus resulting from the rapid pa…

cs.AI2021

Counterfactual Planning in AGI Systems

Koen Holtman

We present counterfactual planning as a design approach for creating a range of safety mechanisms that can be applied in hypothetical future AI systems which have Artificial Genera…

cs.AI2020

AGI Agent Safety by Iteratively Improving the Utility Function

Koen Holtman

While it is still unclear if agents with Artificial General Intelligence (AGI) could ever be built, we can already use mathematical models to investigate potential safety systems f…

cs.AI2021

Demanding and Designing Aligned Cognitive Architectures

Koen Holtman

With AI systems becoming more powerful and pervasive, there is increasing debate about keeping their actions aligned with the broader goals and needs of humanity. This multi-discip…

cs.AI2020

Corrigibility with Utility Preservation

Koen Holtman

Corrigibility is a safety property for artificially intelligent agents. A corrigible agent will not resist attempts by authorized parties to alter the goals and constraints that we…

cs.CY2012

Appcessory Economics: Enabling loosely coupled hardware / software innovation

Koen Holtman

An appcessory (app + accessory) is a smart phone accessory that is combined with a specially written app to perform a useful function. An example is a toy helicopter controlled by…