7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.AI2023
Understanding and Controlling a Maze-Solving Policy Network
Ulisse Mini, Peli Grietzer, Mrinank Sharma +3
To understand the goals and goal representations of AI systems, we carefully study a pretrained reinforcement learning policy that solves mazes by navigating to a range of target s…
cs.CL2023★ 7 cited
Steering Language Models With Activation Engineering
Alexander Matt Turner, Lisa Thiergart, Gavin Leech +4
Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capab…