1 citations · 1 across the 2 of their papers we have counts for
3 papers
Extinction Risks from AI: Invisible to Science?
Vojtech Kovarik, Christian van Merwijk, Ida Mattsson
In an effort to inform the discussion surrounding existential risks from AI, we formulate Extinction-level Goodhart's Law as "Virtually any goal specification, pursued to the extre…
A Complete Criterion for Value of Information in Soluble Influence Diagrams
Chris van Merwijk, Ryan Carey, Tom Everitt
Influence diagrams have recently been used to analyse the safety and fairness properties of AI systems. A key building block for this analysis is a graphical criterion for value of…
Risks from Learned Optimization in Advanced Machine Learning Systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik +2
We analyze the type of learned optimization that occurs when a learned model (such as a neural network) is itself an optimizer - a situation we refer to as mesa-optimization, a neo…