papers

Publications (19)

q-bio.NC2025

The Geometry of Concepts: Sparse Autoencoder Feature Structure

Yuxiao Li, Eric J. Michaud, David D. Baek +3

Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that thi…

cs.LG2026

Building Production-Ready Probes For Gemini

János Kramár, Joshua Engels, Zheng Wang +4

Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…

cs.LG2025

Efficient Dictionary Learning with Switch Sparse Autoencoders

Anish Mudide, Joshua Engels, Eric J. Michaud +2

Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features…

cs.LG2025

Low-Rank Adapting Models for Sparse Autoencoders

Matthew Chen, Joshua Engels, Max Tegmark

Sparse autoencoders (SAEs) decompose language model representations into a sparse set of linear latent vectors. Recent works have improved SAEs using language model gradients, but…

cs.CL2025

Simple Mechanistic Explanations for Out-Of-Context Reasoning

Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2

Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…

cs.LG2026

How Transparent is DiffusionGemma?

Joshua Engels, Callum McDougall, Bilal Chughtai +11

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…