7 papers
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Reef Menaged, Gili Lior, Shauli Ravfogel +2
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
Or Shafran, Shaked Ronen, Omri Fahn +3
Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for in…
Preserving Task-Relevant Information Under Linear Concept Removal
Floris Holstege, Shauli Ravfogel, Bram Wouters
Modern neural networks often encode unwanted concepts alongside task-relevant information, leading to fairness and interpretability concerns. Existing post-hoc approaches can remov…
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
Hung-Ting Chen, Xiang Liu, Shauli Ravfogel +1
Most text retrievers generate \emph{one} query vector to retrieve relevant documents. Yet, the conditional distribution of relevant documents for the query may be multimodal, e.g.,…
Emergence of Linear Truth Encodings in Language Models
Shauli Ravfogel, Gilad Yehudai, Tal Linzen +2
Recent probing studies reveal that large language models exhibit linear subspaces that separate true from false statements, yet the mechanism behind their emergence is unclear. We…
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
Yu Fan, Yang Tian, Shauli Ravfogel +3
Embedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes lik…