299 citations · 478 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2011★ 119 cited
Efficient Optimal Learning for Contextual Bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale +4
We address the problem of learning in an online setting where the learner repeatedly observes features, selects among a set of actions, and receives reward for the action taken. We…
cs.LG2011★ 299 cited
Doubly Robust Policy Evaluation and Learning
Miroslav Dudik, John Langford, Lihong Li
We study decision making in environments where the reward is only partially observed, but can be modeled as a function of an action and an observed context. This setting, known as…