2 papers
cs.LG2026
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
Andreas Haupt, Justin Hartenstein, Anka Reuel +2
AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has received far less attention:…
cs.LG2025
PrivATE: Differentially Private Confidence Intervals for Average Treatment Effects
Maresa Schröder, Justin Hartenstein, Stefan Feuerriegel
The average treatment effect (ATE) is widely used to evaluate the effectiveness of drugs and other medical interventions. In safety-critical applications like medicine, reliable in…