1 paper
Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such a…