◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Max Nadeau

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.CL1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2023

Benchmarks for Detecting Measurement Tampering

Fabien Roger, Ryan Greenblatt, Max Nadeau +2

When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…

cs.CL2023

Circuit Breaking: Removing Model Behaviors with Targeted Ablation

Maximilian Li, Xander Davies, Max Nadeau

Language models often exhibit behaviors that improve performance on a pre-training objective but harm performance on downstream tasks. We propose a novel approach to removing undes…

cs.AI2023

Discovering Variable Binding Circuitry with Desiderata

Xander Davies, Max Nadeau, Nikhil Prakash +2

Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-outp…

cs.AI2023

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies, Claudia Shi +29

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.