4 papers
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Orga…
Covariate-Adjusted Response-Adaptive Design with Delayed Outcomes
Xinwei Ma, Jingshen Wang, Waverly Wei
Covariate-adjusted response-adaptive (CARA) designs have gained widespread adoption for their clear benefits in enhancing experimental efficiency and participant welfare. These des…
SLOACI: Surrogate-Leveraged Online Adaptive Causal Inference
Yingying Fan, Zihan Wang, Waverly Wei
Adaptive experimental designs have gained increasing attention across a range of domains. In this paper, we propose a new methodological framework, surrogate-leveraged online adapt…
Can Language Models Boost the Power of Randomized Experiments Without Statistical Bias?
Xinrui Ruan, Xinwei Ma, Yingfei Wang +2
Randomized controlled trials (RCTs) are widely adopted for causal inference, yet cost and sample-size constraints limit power. We introduce CALM (Causal Analysis leveraging Languag…