2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1
Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…