◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

David Manheim

1Day Sooner

4 papers hereh-index 10541 citations37 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • first author1
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.AI3
  • cs.MA1
affiliations
  • 1Day Sooner
Homepage
same name
  • David Manheim — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.AIShow all

3 papers · 1 filter

cs.AI2022

Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety

Issa Rice, David Manheim

Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of d…

cs.AI2018

Oversight of Unsafe Systems via Dynamic Safety Envelopes

David Manheim

This paper reviews the reasons that Human-in-the-Loop is both critical for preventing widely-understood failure modes for machine learning, and not a practical solution. Following…

cs.AI2018

Categorizing Variants of Goodhart's Law

David Manheim, Scott Garrabrant

There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an exte…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.