1 paper · 1 filter
Zhongren Chen, Siyu Chen, Zhengling Qi +2
We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest I^±-quantile for some I^±∈(0,1). We focus on the offlin…