◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Michael Sellitto

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.CR1
  • cs.CY1

identity via Semantic Scholar / OpenAlex

most citedThe Capacity for Moral Self-Correction in Large Language Models

53 citations · 105 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CR2024★ 39 cited

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Evan Hubinger, Carson Denison, Jesse Mu +36

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…

cs.CY2023★ 13 cited

Confidence-Building Measures for Artificial Intelligence: Workshop Proceedings

Sarah Shoker, Andrew Reddie, Sarah Barrington +20

Foundation models could eventually introduce several pathways for undermining state security: accidents, inadvertent escalation, unintentional conflict, the proliferation of weapon…

cs.CL2023★ 53 cited

The Capacity for Moral Self-Correction in Large Language Models

Deep Ganguli, Amanda Askell, Nicholas Schiefer +46

We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.