3 citations · 3 across the 3 of their papers we have counts for
3 papers
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…
Inducing Human-like Biases in Moral Reasoning Language Models
Artem Karpov, Seong Hah Cho, Austin Meek +3
In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same…
Understanding and Controlling a Maze-Solving Policy Network
Ulisse Mini, Peli Grietzer, Mrinank Sharma +3
To understand the goals and goal representations of AI systems, we carefully study a pretrained reinforcement learning policy that solves mazes by navigating to a range of target s…