1 paper
Pedro R. A. S. Bassi, Wenxuan Li, Yucheng Tang +50
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified…