papers

Publications (6)

cs.LG2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

Adam Karvonen, Benjamin Wright, Can Rager +6

What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representati…

cs.AI2023

Optimal Policies Tend to Seek Power

Alexander Matt Turner, Logan Smith, Rohin Shah +2

Some researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of their objectives. Other researchers p…

cs.CY2022

Researching Alignment Research: Unsupervised Analysis

Jan H. Kirchner, Logan Smith, Jacques Thibodeau +2

AI alignment research is the field of study dedicated to ensuring that artificial intelligence (AI) benefits humans. As machine intelligence gets more advanced, this research is be…

math.CO2017

Connected power domination in graphs

Boris Brimkov, Derek Mikesell, Logan Smith

The study of power domination in graphs arises from the problem of placing a minimum number of measurement devices in an electrical network while monitoring the entire network. A p…

cs.LG2025

Eliciting Latent Predictions from Transformers with the Tuned Lens

Nora Belrose, Igor Ostrovsky, Lev McKinney +5

We analyze transformers from the perspective of iterative inference, seeking to understand how model predictions are refined layer by layer. To do so, we train an affine probe for…

math.CO2018

Power domination throttling

Boris Brimkov, Joshua Carlson, Illya V. Hicks +2

A power dominating set of a graph is a set that colors every vertex of according to the following rules: in the first timestep, every vertex in be…