◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Fabien Roger

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

most citedAlignment faking in large language models

24 citations · 26 across the 3 of their papers we have counts for

collaborators

3 papers

cs.AI2024★ 24 cited

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

cs.LG2023★ 2 cited

Preventing Language Models From Hiding Their Reasoning

Fabien Roger, Ryan Greenblatt

Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to moni…

cs.LG2023

Benchmarks for Detecting Measurement Tampering

Fabien Roger, Ryan Greenblatt, Max Nadeau +2

When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.