3 papers
cs.AI2018
Oversight of Unsafe Systems via Dynamic Safety Envelopes
David Manheim
This paper reviews the reasons that Human-in-the-Loop is both critical for preventing widely-understood failure modes for machine learning, and not a practical solution. Following…
cs.MA2018
Multiparty Dynamics and Failure Modes for Machine Learning and Artificial Intelligence
David Manheim
An important challenge for safety in machine learning and artificial intelligence systems is a~set of related failures involving specification gaming, reward hacking, fragility to…
cs.AI2018
Categorizing Variants of Goodhart's Law
David Manheim, Scott Garrabrant
There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an exte…