1 paper · 1 filter
Avi Parrack, Carlo Leonardo Attubato, Stefan Heimersheim
AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activati…