1 paper · 1 filter
Erik Nordby, Tasha Pais, Aviel Parrack
Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fr…