88 citations · 110 across the 3 of their papers we have counts for
3 papers
cs.GT2012★ 15 cited
Deterministic MDPs with Adversarial Rewards and Bandit Feedback
Raman Arora, Ofer Dekel, Ambuj Tewari
We consider a Markov decision process with deterministic state transition dynamics, adversarially generated rewards that change arbitrarily from round to round, and a bandit feedba…
cs.LG2012★ 88 cited
Online Bandit Learning against an Adaptive Adversary: from Regret to Policy Regret
Raman Arora, Ofer Dekel, Ambuj Tewari
Online learning algorithms are designed to learn even when their input is generated by an adversary. The widely-accepted formal definition of an online algorithm's ability to learn…
cs.LG2010★ 7 cited
Robust Distributed Online Prediction
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir +1
The standard model of online prediction deals with serial processing of inputs by a single processor. However, in large-scale online prediction problems, where inputs arrive at a h…