Publications (19)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
Yuxiao Li, Eric J. Michaud, David D. Baek +3
Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that thi…
Building Production-Ready Probes For Gemini
János Kramár, Joshua Engels, Zheng Wang +4
Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…
Efficient Dictionary Learning with Switch Sparse Autoencoders
Anish Mudide, Joshua Engels, Eric J. Michaud +2
Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features…
Low-Rank Adapting Models for Sparse Autoencoders
Matthew Chen, Joshua Engels, Max Tegmark
Sparse autoencoders (SAEs) decompose language model representations into a sparse set of linear latent vectors. Recent works have improved SAEs using language model gradients, but…
Simple Mechanistic Explanations for Out-Of-Context Reasoning
Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2
Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…
How Transparent is DiffusionGemma?
Joshua Engels, Callum McDougall, Bilal Chughtai +11
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…