collaborators

7 papers

cs.CL2026

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

Reef Menaged, Gili Lior, Shauli Ravfogel +2

We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…

cs.CL2026

From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

Or Shafran, Shaked Ronen, Omri Fahn +3

Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for in…

cs.LG2025

Preserving Task-Relevant Information Under Linear Concept Removal

Floris Holstege, Shauli Ravfogel, Bram Wouters

Modern neural networks often encode unwanted concepts alongside task-relevant information, leading to fairness and interpretability concerns. Existing post-hoc approaches can remov…

cs.CL2025

Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval

Hung-Ting Chen, Xiang Liu, Shauli Ravfogel +1

Most text retrievers generate \emph{one} query vector to retrieve relevant documents. Yet, the conditional distribution of relevant documents for the query may be multimodal, e.g.,…

cs.CL2025

Emergence of Linear Truth Encodings in Language Models

Shauli Ravfogel, Gilad Yehudai, Tal Linzen +2

Recent probing studies reveal that large language models exhibit linear subspaces that separate true from false statements, yet the mechanism behind their emergence is unclear. We…

cs.CL2025

The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure

Yu Fan, Yang Tian, Shauli Ravfogel +3

Embedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes lik…