8 papers
Leaderboard Incentives: Model Rankings under Strategic Post-Training
Yatong Chen, Guanhua Zhang, Moritz Hardt
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmax…
Policy Design in Long-Run Welfare Dynamics
Jiduan Wu, Rediet Abebe, Moritz Hardt +1
Improving social welfare is a complex challenge requiring policymakers to optimize objectives across multiple time horizons. Evaluating the impact of such policies presents a funda…
Good Allocations from Bad Estimates
SÃlvia Casacuberta, Moritz Hardt
Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects…
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefor…
Train-before-Test Harmonizes Language Model Rankings
Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt
Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model…
How Benchmark Prediction from Fewer Data Misses the Mark
Guanhua Zhang, Florian E. Dorner, Moritz Hardt
Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets. Benchmark prediction (also cal…