2 papers
cs.LG2026
Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling
Erik Nordby, Tasha Pais, Aviel Parrack
Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fr…
cs.AI2026
Benchmarking Deception Probes via Black-to-White Performance Boosts
Avi Parrack, Carlo Leonardo Attubato, Stefan Heimersheim
AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activati…