3 papers
cs.LG2024
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefor…
cs.GT2024
Causal Inference from Competing Treatments
Ana-Andreea Stoica, Vivian Y. Nastl, Moritz Hardt
Many applications of RCTs involve the presence of multiple treatment administrators -- from field experiments to online advertising -- that compete for the subjects' attention. In…
cs.LG2024
Do causal predictors generalize better to new domains?
Vivian Y. Nastl, Moritz Hardt
We study how well machine learning models trained on causal features generalize across domains. We consider 16 prediction tasks on tabular datasets covering applications in health,…