10 papers
Bi-Orthogonal Factor Decomposition for Vision Transformers
Fenil R. Doshi, Thomas Fel, Talia Konkle +1
Self-attention is the central computational primitive of Vision Transformers, yet we lack a principled understanding of what information attention mechanisms exchange between token…
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
Gaia Grosso, Sai Sumedh R. Hindupur, Thomas Fel +3
Modern artificial intelligence has revolutionized our ability to extract rich and versatile data representations across scientific disciplines. Yet, the statistical properties of t…
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
Fenil R. Doshi, Thomas Fel, Talia Konkle +1
Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primarily harvest local texture cues, yi…
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
Matyas Bohacek, Thomas Fel, Maneesh Agrawala +1
Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands…
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
Valérie Costa, Thomas Fel, Ekdeep Singh Lubana +2
Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural…