2 papers
cs.CY2026
RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies
Patricia Paskov, Kevin Wei, Shen Zhou Hong +7
Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform fro…
cs.AI2025
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
Kevin L. Wei, Patricia Paskov, Sunishchal Dev +6
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI pe…