5 papers
Structure Before Collapse: Transient semantic geometry in next-token prediction
Yize Zhao, Isabel Papadimitriou, Christos Thrampoulidis
Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the out…
Vocabulary embeddings organize linguistic structure early in language model training
Isabel Papadimitriou, Jacob Prince
Large language models (LLMs) work by manipulating the geometry of input embedding vectors over multiple layers. Here, we ask: how are the input vocabulary representations of langua…
Interpreting the linear structure of vision-language model embedding spaces
Isabel Papadimitriou, Huangyuan Su, Thomas Fel +2
Vision-language models encode images and text in a joint space, minimizing the distance between corresponding image and text pairs. How are language and images organized in this jo…
Using Shapley interactions to understand how models use structure
Divyansh Singhvi, Diganta Misra, Andrej Erkelens +3
Language is an intricately structured system, and a key goal of NLP interpretability is to provide methodological insights for understanding how language models represent this stru…
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince +7
Sparse Autoencoders (SAEs) have emerged as a powerful framework for machine learning interpretability, enabling the unsupervised decomposition of model representations into a dicti…