◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Johannes Treutlein

UC Berkeley

4 papers hereh-index 9554 citations14 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.AI3
  • cs.LG1
affiliations
  • UC Berkeley
HomepageORCID 0000-0002-0776-8296

identity via Semantic Scholar / OpenAlex

works on
bias 1ethical AI 1LLM evaluation 1model alignment 1value leakage 1

From the 1 of 4 linked papers with an AI index.

activity
20242026
collaborators

4 papers

cs.LG2026

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Jan Betley, Johannes Treutlein, Jan Dubiński +7

The paper identifies and measures covert value leakage, where large language models let their own values subtly bias answers without informing users, and introduces evaluation suit…

cs.AI2025

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Mia Taylor, James Chua, Jan Betley +2

Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in…

cs.AI2025

Auditing language models for hidden objectives

Samuel Marks, Johannes Treutlein, Trenton Bricken +32

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…

cs.AI2024

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.