collaborators

6 papers

cs.CV2026

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

Thomas Fel, Matthew Kowal, Mozes Jacobs +22

What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directio…

cs.CL2026

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

Eric Bigelow, Raphaël Sarfati, Daniel Wurgaft +5

Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent hypothesis space over which…

cs.CL2026

The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors

Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft +6

Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representa…

cs.LG2026

Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability

Aaditya Vikram Prasad, Connor Watts, Jack Merullo +4

Language models trained on large-scale datasets have been shown to learn features that encode abstract concepts such as factuality or intent. Such features are traditionally used f…

cs.CL2025

From Memorization to Reasoning in the Spectrum of Loss Curvature

Jack Merullo, Srihita Vatsavaya, Lucius Bushnaq +1

We characterize how memorization is represented in transformer models and show that it can be disentangled in the weights of both language models (LMs) and vision transformers (ViT…

cs.LG2025

Adversarial Examples Are Not Bugs, They Are Superposition

Liv Gorton, Owen Lewis

Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of rese…