collaborators

37 papers

cs.CL2026

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis

Amit Elhelo, Amir Globerson, Mor Geva

Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks, motivating the view of their parameters as a knowledge base. An important prop…

cs.CL2026

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

Clara Haya Suslik, Or Shafran, Mor Geva

As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety and compliance. Prominent met…

cs.CL2026

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

Yoav Gur-Arieh, Mor Geva, Atticus Geiger

A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might represent "Ann loves pie" by binding…

cs.CL2026

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

Yoav Gur-Arieh, Ana Marasović, Mor Geva

Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithful…

cs.CL2026

Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models

Itay Yona, Dan Barzilay, Michael Karasik +1

How do language models retrieve entity-specific facts from their parameters? We investigate this question by searching for sparse, entity-selective MLP neurons - which we call enti…

cs.CV2026

Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex

Idan Daniel Grosbard, Mor Geva, Galit Yovel

A central goal in understanding human vision is to uncover the visual features that drive neuronal activity. A growing body of work has used artificial neural networks as encoding…