Showing cs.AIShow all
3 papers · 1 filter
cs.AI2022
Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety
Issa Rice, David Manheim
Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of d…
cs.AI2018
Oversight of Unsafe Systems via Dynamic Safety Envelopes
David Manheim
This paper reviews the reasons that Human-in-the-Loop is both critical for preventing widely-understood failure modes for machine learning, and not a practical solution. Following…
cs.AI2018
Categorizing Variants of Goodhart's Law
David Manheim, Scott Garrabrant
There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an exte…