3 papers
cs.LG2026
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Lei Shi, Anlan Zhang, Rita Lyu +6
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges…
cs.AI2026
Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking
Zhengmian Hu, Lei Shi, Ritwik Sinha +2
Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and oft…
stat.ME2025
Leveraging semantic similarity for experimentation with AI-generated treatments
Lei Shi, David Arbour, Raghavendra Addanki +2
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main me…