Publications (6)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Adam Karvonen, Benjamin Wright, Can Rager +6
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representati…
Optimal Policies Tend to Seek Power
Alexander Matt Turner, Logan Smith, Rohin Shah +2
Some researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of their objectives. Other researchers p…
Researching Alignment Research: Unsupervised Analysis
Jan H. Kirchner, Logan Smith, Jacques Thibodeau +2
AI alignment research is the field of study dedicated to ensuring that artificial intelligence (AI) benefits humans. As machine intelligence gets more advanced, this research is be…
Connected power domination in graphs
Boris Brimkov, Derek Mikesell, Logan Smith
The study of power domination in graphs arises from the problem of placing a minimum number of measurement devices in an electrical network while monitoring the entire network. A p…
Eliciting Latent Predictions from Transformers with the Tuned Lens
Nora Belrose, Igor Ostrovsky, Lev McKinney +5
We analyze transformers from the perspective of iterative inference, seeking to understand how model predictions are refined layer by layer. To do so, we train an affine probe for…
Power domination throttling
Boris Brimkov, Joshua Carlson, Illya V. Hicks +2
A power dominating set of a graph is a set that colors every vertex of according to the following rules: in the first timestep, every vertex in be…