2 papers
stat.ME2026
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Orga…
stat.ME2025
Can Language Models Boost the Power of Randomized Experiments Without Statistical Bias?
Xinrui Ruan, Xinwei Ma, Yingfei Wang +2
Randomized controlled trials (RCTs) are widely adopted for causal inference, yet cost and sample-size constraints limit power. We introduce CALM (Causal Analysis leveraging Languag…