37 papers
LMs as Task-Specific Knowledge Bases: An Interpretability Analysis
Amit Elhelo, Amir Globerson, Mor Geva
Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks, motivating the view of their parameters as a knowledge base. An important prop…
Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings
Clara Haya Suslik, Or Shafran, Mor Geva
As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety and compliance. Prominent met…
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
Yoav Gur-Arieh, Mor Geva, Atticus Geiger
A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might represent "Ann loves pie" by binding…
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
Yoav Gur-Arieh, Ana MarasoviÄ, Mor Geva
Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithful…
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
Itay Yona, Dan Barzilay, Michael Karasik +1
How do language models retrieve entity-specific facts from their parameters? We investigate this question by searching for sparse, entity-selective MLP neurons - which we call enti…
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
Idan Daniel Grosbard, Mor Geva, Galit Yovel
A central goal in understanding human vision is to uncover the visual features that drive neuronal activity. A growing body of work has used artificial neural networks as encoding…