3 papers
cs.CL2026
Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings
Clara Haya Suslik, Or Shafran, Mor Geva
As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety and compliance. Prominent met…
cs.CL2026
Constructing Interpretable Features from Compositional Neuron Groups
Or Shafran, Atticus Geiger, Mor Geva
A central goal for mechanistic interpretability has been to identify the right units of analysis in large language models (LLMs) that causally explain their outputs. While early wo…
cs.CL2026
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
Or Shafran, Shaked Ronen, Omri Fahn +3
Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for in…