activity
20242026
collaborators

16 papers

cs.LG2026

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath

Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify mod…

cs.CL2026

From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation

Aviya Maimon, Amir DN Cohen, Gal Vishne +2

Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this compariso…

cs.AI2026

Can LLMs Introspect? A Reality Check

Shashwat Singh, Tal Linzen, Shauli Ravfogel

Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, w…

cs.CL2026

Geometric Factual Recall in Transformers

Shauli Ravfogel, Gilad Yehudai, Joan Bruna +1

How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of embeddings, requiring parameter…

cs.CL2026

RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context

Jackson Petty, Michael Y. Hu, Wentao Wang +3

Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For m…

cs.LG2026

The Truthfulness Spectrum Hypothesis

Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte +1

Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness…