Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Michael Hardy, Anka Reuel, Lijin Zhang +6
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…
cs.AI2026
Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach
Jodi M. Casabianca, Maggie Beiting-Parrish
Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psyc…