◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Brian Christian

3 papers hereh-index 344 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CLShow all

2 papers · 1 filter

cs.CL2026

Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models

Brian Christian, Matan Mazor

Fair decisions require ignoring irrelevant, potentially biasing, information. To achieve this, decision-makers need to approximate what decision they would have made had they not k…

cs.CL2025

Reward Model Interpretability via Optimal and Pessimal Tokens

Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2

Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.