1 paper
Michael Hardy, Anka Reuel, Lijin Zhang +6
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…