5 papers
Comparing Linear Probes with Mahalanobis Cosine Similarity
Zhuofan Josh Ying, Peter Hase, Nikolaus Kriegeskorte
Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis cosine similarity (MCS) between two directions, which reweights…
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
Peter Hase, Christopher Potts
Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with CoT faithfulness severely limit w…
The Truthfulness Spectrum Hypothesis
Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte +1
Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness…
Unsupervised Elicitation of Language Models
Jiaxin Wen, Zachary Ankner, Arushi Somani +10
To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabili…
Reasoning Models Don't Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan +12
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…