collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2025

Investigating Representation Universality: Case Study on Genealogical Representations

David D. Baek, Yuxiao Li, Max Tegmark

Motivated by interpretability and reliability, we investigate whether large language models (LLMs) deploy universal geometric structures to encode discrete, graph-structured knowle…

cs.LG2025

Dense SAE Latents Are Features, Not Bugs

Xiaoqing Sun, Alessandro Stolfo, Joshua Engels +4

Sparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that…

cs.LG2025

Efficient Dictionary Learning with Switch Sparse Autoencoders

Anish Mudide, Joshua Engels, Eric J. Michaud +2

Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features…

cs.LG2025

Low-Rank Adapting Models for Sparse Autoencoders

Matthew Chen, Joshua Engels, Max Tegmark

Sparse autoencoders (SAEs) decompose language model representations into a sparse set of linear latent vectors. Recent works have improved SAEs using language model gradients, but…

cs.LG2025

Towards Understanding Distilled Reasoning Models: A Representational Approach

David D. Baek, Max Tegmark

In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs). To explore this, we train a crosscoder on Qwen-s…

cs.LG2025

Decomposing The Dark Matter of Sparse Autoencoders

Joshua Engels, Logan Riggs, Max Tegmark

Sparse autoencoders (SAEs) are a promising technique for decomposing language model activations into interpretable linear features. However, current SAEs fall short of completely e…