◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

David Manheim

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • first author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.MA1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2018

Oversight of Unsafe Systems via Dynamic Safety Envelopes

David Manheim

This paper reviews the reasons that Human-in-the-Loop is both critical for preventing widely-understood failure modes for machine learning, and not a practical solution. Following…

cs.MA2018

Multiparty Dynamics and Failure Modes for Machine Learning and Artificial Intelligence

David Manheim

An important challenge for safety in machine learning and artificial intelligence systems is a~set of related failures involving specification gaming, reward hacking, fragility to…

cs.AI2018

Categorizing Variants of Goodhart's Law

David Manheim, Scott Garrabrant

There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an exte…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.