activity
20242026
collaborators

6 papers

cs.AI2026

Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs

Jonathan Cook, Silvia Sapora, Arash Ahmadian +4

Large language models (LLMs) are typically trained to acquire behaviours from demonstrations or experience, yet much of their training data is declarative: instructions, rules, and…

cs.LG2025

Ctrl-Z: Controlling AI Agents via Resampling

Aryan Bhatt, Cody Rushing, Adam Kaufman +5

Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first contr…

cs.AI2025

Auditing language models for hidden objectives

Samuel Marks, Johannes Treutlein, Trenton Bricken +32

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…

cs.MA2025

Factorio Learning Environment

Jack Hopkins, Mart Bakler, Akbir Khan

Large Language Models (LLMs) are rapidly saturating existing benchmarks, necessitating new open-ended evaluations. We introduce the Factorio Learning Environment (FLE), based on th…

cs.AI2024

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

cs.CL2024

Language Models Learn to Mislead Humans via RLHF

Jiaxin Wen, Ruiqi Zhong, Akbir Khan +6

Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this p…