Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
Jonathan Cook, Silvia Sapora, Arash Ahmadian +4
Large language models (LLMs) are typically trained to acquire behaviours from demonstrations or experience, yet much of their training data is declarative: instructions, rules, and…
cs.AI2025
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…
cs.AI2024
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright +17
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…