3 papers
stat.ME2026
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Orga…
stat.ME2025
Covariate-Adjusted Response-Adaptive Design with Delayed Outcomes
Xinwei Ma, Jingshen Wang, Waverly Wei
Covariate-adjusted response-adaptive (CARA) designs have gained widespread adoption for their clear benefits in enhancing experimental efficiency and participant welfare. These des…
stat.ME2025
Can Language Models Boost the Power of Randomized Experiments Without Statistical Bias?
Xinrui Ruan, Xinwei Ma, Yingfei Wang +2
Randomized controlled trials (RCTs) are widely adopted for causal inference, yet cost and sample-size constraints limit power. We introduce CALM (Causal Analysis leveraging Languag…