◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ryan Greenblatt

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.CR1

identity via Semantic Scholar / OpenAlex

most citedSleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

39 citations · 41 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CR2024★ 39 cited

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Evan Hubinger, Carson Denison, Jesse Mu +36

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…

cs.LG2023★ 2 cited

Preventing Language Models From Hiding Their Reasoning

Fabien Roger, Ryan Greenblatt

Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to moni…

cs.LG2023

Benchmarks for Detecting Measurement Tampering

Fabien Roger, Ryan Greenblatt, Max Nadeau +2

When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.