collaborators

8 papers

cs.GT2026

Leaderboard Incentives: Model Rankings under Strategic Post-Training

Yatong Chen, Guanhua Zhang, Moritz Hardt

Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmax…

cs.CY2026

Policy Design in Long-Run Welfare Dynamics

Jiduan Wu, Rediet Abebe, Moritz Hardt +1

Improving social welfare is a complex challenge requiring policymakers to optimize objectives across multiple time horizons. Evaluating the impact of such policies presents a funda…

cs.LG2026

Good Allocations from Bad Estimates

Sílvia Casacuberta, Moritz Hardt

Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects…

cs.LG2026

Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data

Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt

High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefor…

cs.LG2025

Train-before-Test Harmonizes Language Model Rankings

Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt

Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model…

cs.LG2025

How Benchmark Prediction from Fewer Data Misses the Mark

Guanhua Zhang, Florian E. Dorner, Moritz Hardt

Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets. Benchmark prediction (also cal…