Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Efficient Evaluation of LLM Performance with Statistical Guarantees
Skyler Wu, Yash Nair, Emmanuel J. Candès
Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query…
stat.ML2025
Missing Data Multiple Imputation for Tabular Q-Learning in Online RL
Kyla Chasalow, Skyler Wu, Susan Murphy
Missing data in online reinforcement learning (RL) poses challenges compared to missing data in standard tabular data or in offline policy learning. The need to impute and act at e…