5 citations · 7 across the 6 of their papers we have counts for
1 paper · 1 filter
Jiashuo Jiang, Yinyu Ye, Yiming Zong
Learning the optimal policy for Markov decision process problems (MDPs) from samples is a fundamental problem in online and data-driven decision-making. Function approximations are…