activity
20242026
most citedInto the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry

4 citations · 4 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

24 papers · 1 filter

cs.LG2026

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

Leon Bergen, Usha Bhalla, Sidharth Baskaran +14

Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. Th…

cs.LG2026

From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?

Aaron Mueller, Andrew Lee, Shruti Joshi +3

A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…

cs.LG2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Jing Huang, Daniel Wurgaft, Rachit Bansal +6

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger mo…

cs.LG2026

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

Daniel Wurgaft, Can Rager, Matthew Kowal +13

Neural representations carry rich geometric structure; but does that structure causally shape behavior? To address this question, we intervene along paths through activation space…

cs.LG2026

Do Sparse Autoencoders Capture Concept Manifolds?

Usha Bhalla, Thomas Fel, Can Rager +9

Sparse autoencoders (SAEs) are widely used to extract interpretable features from neural network representations, often under the implicit assumption that concepts correspond to in…

cs.LG2026

Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

Eric Bigelow, Daniel Wurgaft, YingQiao Wang +4

Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been…