1 paper · 1 filter
Joe Needham, Giles Edkins, Govind Pimpale +2
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior durin…