7 citations · 24 across the 7 of their papers we have counts for
Showing 2016Show all
2 papers · 1 filter
cs.AI2016★ 7 cited
Optimizing Quantiles in Preference-based Markov Decision Processes
Hugo Gilbert, Paul Weng, Yan Xu
In the Markov decision process model, policies are usually evaluated by expected cumulative rewards. As this decision criterion is not always suitable, we propose in this paper an…
cs.LG2016★ 6 cited
Quantile Reinforcement Learning
Hugo Gilbert, Paul Weng
In reinforcement learning, the standard criterion to evaluate policies in a state is the expectation of (discounted) sum of rewards. However, this criterion may not always be suita…