2 papers
cs.LG2026
Comparing Linear Probes with Mahalanobis Cosine Similarity
Zhuofan Josh Ying, Peter Hase, Nikolaus Kriegeskorte
Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis cosine similarity (MCS) between two directions, which reweights…
cs.LG2026
The Truthfulness Spectrum Hypothesis
Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte +1
Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness…