activity
20182021
most citedBeyond Value-Function Gaps: Improved Instance-Dependent Regret Bounds for Episodic Reinforcement Learning

5 citations · 5 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG2021

The Pareto Frontier of model selection for general Contextual Bandits

Teodor V. Marinov, Julian Zimmert

Recent progress in model selection raises the question of the fundamental limits of these techniques. Under specific scrutiny has been model selection for general contextual bandit…

cs.LG20215 cited

Beyond Value-Function Gaps: Improved Instance-Dependent Regret Bounds for Episodic Reinforcement Learning

Christoph Dann, Teodor V. Marinov, Mehryar Mohri +1

We provide improved gap-dependent regret bounds for reinforcement learning in finite episodic Markov decision processes. Compared to prior work, our bounds depend on alternative de…

cs.LG2020

Corralling Stochastic Bandit Algorithms

Raman Arora, Teodor V. Marinov, Mehryar Mohri

We study the problem of corralling stochastic bandit algorithms, that is combining multiple bandit algorithms designed for a stochastic environment, with the goal of devising a cor…

cs.LG2020

Private Stochastic Convex Optimization: Efficient Algorithms for Non-smooth Objectives

Raman Arora, Teodor V. Marinov, Enayat Ullah

In this paper, we revisit the problem of private stochastic convex optimization. We propose an algorithm based on noisy mirror descent, which achieves optimal rates both in terms o…

cs.LG2019

Bandits with Feedback Graphs and Switching Costs

Raman Arora, Teodor V. Marinov, Mehryar Mohri

We study the adversarial multi-armed bandit problem where partial observations are available and where, in addition to the loss incurred for each action, a \emph{switching cost} is…

cs.LG2018

Policy Regret in Repeated Games

Raman Arora, Michael Dinitz, Teodor V. Marinov +1

The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such…