2 papers
cs.CL2026
Probe Generalization as Subspace Selection for OOD Deception Detection
Daniel Yoo, Adrians Skapars
Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generaliza…
cs.AI2025
Caught in the Act: a mechanistic approach to detecting deception
Gerard Boxo, Ryan Socha, Daniel Yoo +1
Sophisticated instrumentation for AI systems might have indicators that signal misalignment from human values, not unlike a "check engine" light in cars. One such indicator of misa…