2 citations · 2 across the 8 of their papers we have counts for
11 papers
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Lei Shi, Anlan Zhang, Rita Lyu +6
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges…
Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems
Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou +27
Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings -…
AI-Assisted Variance Reduction in Randomized Experiments
David Arbour, Eli Ben-Michael, Avi Feller +2
Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent wor…
A Weighting Framework for Clusters as Confounders in Observational Studies
Eli Ben-Michael, Avi Feller, Luke Keele
When units in observational studies are clustered in groups, such as students in schools or patients in hospitals, researchers often address confounding by adjusting for cluster-le…
Forest Kernel Balancing Weights: Outcome-Guided Features for Causal Inference
Andy A. Shen, Eli Ben-Michael, Avi Feller +2
While balancing covariates between groups is central for observational causal inference, selecting which features to balance remains a challenging problem. Kernel balancing is a pr…
Leveraging semantic similarity for experimentation with AI-generated treatments
Lei Shi, David Arbour, Raghavendra Addanki +2
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main me…