7 papers
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Lei Shi, Anlan Zhang, Rita Lyu +6
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges…
AI-Assisted Variance Reduction in Randomized Experiments
David Arbour, Eli Ben-Michael, Avi Feller +2
Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent wor…
Regularizing Extrapolation in Causal Inference
David Arbour, Harsh Parikh, Bijan Niknam +3
Many common estimators in machine learning and causal inference are linear smoothers, where the prediction is a weighted average of the training outcomes. Some estimators, such as…
Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking
Zhengmian Hu, Lei Shi, Ritwik Sinha +2
Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and oft…
Leveraging semantic similarity for experimentation with AI-generated treatments
Lei Shi, David Arbour, Raghavendra Addanki +2
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main me…
Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation
Zhenghao Zeng, David Arbour, Avi Feller +3
Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of i…