1 paper · 1 filter
Yang Xu, Washim Uddin Mondal, Vaneet Aggarwal
We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asympto…