actor-critic 2average reward 1average-reward MDP 1constrained mdp 1distributional robustness 1multilevel monte carlo 1neural networks 1q-learning 1robust reinforcement learning 1
From the 2 of 37 linked papers with an AI index.
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
Yang Xu, Jiefu Zhang, Haixiang Sun +3
Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…
stat.ML2026
Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients
Yang Xu, Vaneet Aggarwal
We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnosti…
stat.ML2025
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
Yang Xu, Washim Uddin Mondal, Vaneet Aggarwal
We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asympto…