26 citations · 59 across the 18 of their papers we have counts for
20 papers · 1 filter
Data-Driven Online Model Selection With Regret Guarantees
Aldo Pacchiano, Christoph Dann, Claudio Gentile
We consider model selection for sequential decision making in stochastic environments with bandit feedback, where a meta-learner has at its disposal a pool of base learners, and de…
A Blackbox Approach to Best of Both Worlds in Bandits and Beyond
Christoph Dann, Chen-Yu Wei, Julian Zimmert
Best-of-both-worlds algorithms for online learning which achieve near-optimal regret in both the adversarial and the stochastic regimes have received growing attention recently. Ex…
Best of Both Worlds Policy Optimization
Christoph Dann, Chen-Yu Wei, Julian Zimmert
Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving regret bounds e…
Pseudonorm Approachability and Applications to Regret Minimization
Christoph Dann, Yishay Mansour, Mehryar Mohri +2
Blackwell's celebrated approachability theory provides a general framework for a variety of learning problems, including regret minimization. However, Blackwell's proof and implici…
Learning in POMDPs is Sample-Efficient with Hindsight Observability
Jonathan N. Lee, Alekh Agarwal, Christoph Dann +1
POMDPs capture a broad class of decision making problems, but hardness results suggest that learning is intractable even in simple settings due to the inherent partial observabilit…
A Unified Algorithm for Stochastic Path Problems
Christoph Dann, Chen-Yu Wei, Julian Zimmert
We study reinforcement learning in stochastic path (SP) problems. The goal in these problems is to maximize the expected sum of rewards until the agent reaches a terminal state. We…