1 paper
Jingyi Li, Peng Wu, Chengchun Shi
Reinforcement learning algorithms are generally designed to maximize the expected return across a population. However, a policy that is optimal on average may be suboptimal for cer…