1 paper · 1 filter
Vojtech Kovarik, Eric Olav Chen, Sami Petersen +2
This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility…