2 papers
cs.CY2025
Preliminary suggestions for rigorous GPAI model evaluations
Patricia Paskov, Michael J. Byun, Kevin Wei +1
This document presents a preliminary compilation of general-purpose AI (GPAI) evaluation practices that may promote internal validity, external validity and reproducibility. It inc…
cs.AI2025
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
Kevin L. Wei, Patricia Paskov, Sunishchal Dev +6
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI pe…