1 paper
Kevin L. Wei, Patricia Paskov, Sunishchal Dev +6
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI pe…