16 papers
What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath
Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify mod…
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
Aviya Maimon, Amir DN Cohen, Gal Vishne +2
Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this compariso…
Can LLMs Introspect? A Reality Check
Shashwat Singh, Tal Linzen, Shauli Ravfogel
Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, w…
Geometric Factual Recall in Transformers
Shauli Ravfogel, Gilad Yehudai, Joan Bruna +1
How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of embeddings, requiring parameter…
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
Jackson Petty, Michael Y. Hu, Wentao Wang +3
Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For m…
The Truthfulness Spectrum Hypothesis
Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte +1
Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness…