6 papers
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Thomas Fel, Matthew Kowal, Mozes Jacobs +22
What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directio…
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
Eric Bigelow, Raphaël Sarfati, Daniel Wurgaft +5
Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent hypothesis space over which…
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft +6
Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representa…
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
Aaditya Vikram Prasad, Connor Watts, Jack Merullo +4
Language models trained on large-scale datasets have been shown to learn features that encode abstract concepts such as factuality or intent. Such features are traditionally used f…
From Memorization to Reasoning in the Spectrum of Loss Curvature
Jack Merullo, Srihita Vatsavaya, Lucius Bushnaq +1
We characterize how memorization is represented in transformer models and show that it can be disentangled in the weights of both language models (LMs) and vision transformers (ViT…
Adversarial Examples Are Not Bugs, They Are Superposition
Liv Gorton, Owen Lewis
Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of rese…