6 citations · 9 across the 3 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2021
Counterfactual Planning in AGI Systems
Koen Holtman
We present counterfactual planning as a design approach for creating a range of safety mechanisms that can be applied in hypothetical future AI systems which have Artificial Genera…
cs.AI2020★ 6 cited
AGI Agent Safety by Iteratively Improving the Utility Function
Koen Holtman
While it is still unclear if agents with Artificial General Intelligence (AGI) could ever be built, we can already use mathematical models to investigate potential safety systems f…
cs.AI2019
Corrigibility with Utility Preservation
Koen Holtman
Corrigibility is a safety property for artificially intelligent agents. A corrigible agent will not resist attempts by authorized parties to alter the goals and constraints that we…