11 papers
Interpreting Brain Responses to Language with Sparse Features from Language Models
Michael A. Lepori, Kendrick Kay, Greta Tuckute
A central goal of cognitive neuroscience is to characterize the features that are represented by human language cortex. Artificial language models (LMs) have emerged as a powerful…
Different types of syntactic agreement recruit the same units within large language models
Daria Kryvosheieva, Andrea de Varda, Evelina Fedorenko +1
Large language models (LLMs) can reliably distinguish grammatical from ungrammatical sentences, but how grammatical knowledge is represented within the models remains an open quest…
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
Taha Binhuraib, Greta Tuckute, Nicholas Blauch
Spatial functional organization is a hallmark of biological brains: neurons are arranged topographically according to their response properties, at multiple scales. In contrast, re…
Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
Shreya Saha, Shurui Li, Greta Tuckute +5
The human language system represents both linguistic forms and meanings, but the abstractness of the meaning representations remains debated. Here, we searched for abstract represe…
From Language to Cognition: How LLMs Outgrow the Human Language Network
Badr AlKhamissi, Greta Tuckute, Yingtian Tang +3
Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language shaping brain-like representati…