4 papers
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Michael Hardy, Anka Reuel, Lijin Zhang +6
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
Jodi M. Casabianca, Daniel F. McCaffrey, Matthew S. Johnson +2
The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more…
Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach
Jodi M. Casabianca, Maggie Beiting-Parrish
Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psyc…
Validity Arguments For Constructed Response Scoring Using Generative Artificial Intelligence Applications
Jodi M. Casabianca, Daniel F. McCaffrey, Matthew S. Johnson +2
The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more…