◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nevan Wichers

3 papers hereh-index 103.7k citations20 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2026

Model Spec Midtraining: Improving How Alignment Training Generalizes

Chloe Li, Nevan Wichers, Sara Price +2

Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…

cs.AI2026

Recontextualization Mitigates Specification Gaming without Modifying the Specification

Ariana Azarbal, Victor Gillioz, Vladimir Ivanov +6

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game"…

cs.LG2025

Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment

Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8

Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.