Showing stat.MEShow all
2 papers · 1 filter
stat.ME2025
Leveraging semantic similarity for experimentation with AI-generated treatments
Lei Shi, David Arbour, Raghavendra Addanki +2
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main me…
stat.ME2025
Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation
Zhenghao Zeng, David Arbour, Avi Feller +3
Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of i…