1 paper · 1 filter
Lorenzo Pacchiardi, Marko Tesic, Lucy G. Cheke +1
The integrity of AI benchmarks is fundamental to accurately assess the capabilities of AI systems. The internal validity of these benchmarks - i.e., making sure they are free from…