1 paper
Amr Moustafa, Max Feser, Florian Mai
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain…