5 citations · 5 across the 2 of their papers we have counts for
7 papers
The Pareto Frontier of model selection for general Contextual Bandits
Teodor V. Marinov, Julian Zimmert
Recent progress in model selection raises the question of the fundamental limits of these techniques. Under specific scrutiny has been model selection for general contextual bandit…
Beyond Value-Function Gaps: Improved Instance-Dependent Regret Bounds for Episodic Reinforcement Learning
Christoph Dann, Teodor V. Marinov, Mehryar Mohri +1
We provide improved gap-dependent regret bounds for reinforcement learning in finite episodic Markov decision processes. Compared to prior work, our bounds depend on alternative de…
Corralling Stochastic Bandit Algorithms
Raman Arora, Teodor V. Marinov, Mehryar Mohri
We study the problem of corralling stochastic bandit algorithms, that is combining multiple bandit algorithms designed for a stochastic environment, with the goal of devising a cor…
Private Stochastic Convex Optimization: Efficient Algorithms for Non-smooth Objectives
Raman Arora, Teodor V. Marinov, Enayat Ullah
In this paper, we revisit the problem of private stochastic convex optimization. We propose an algorithm based on noisy mirror descent, which achieves optimal rates both in terms o…
Bandits with Feedback Graphs and Switching Costs
Raman Arora, Teodor V. Marinov, Mehryar Mohri
We study the adversarial multi-armed bandit problem where partial observations are available and where, in addition to the loss incurred for each action, a \emph{switching cost} is…
Policy Regret in Repeated Games
Raman Arora, Michael Dinitz, Teodor V. Marinov +1
The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such…