activity
20152024
most citedRegret Bound Balancing and Elimination for Model Selection in Bandits and RL

12 citations · 90 across the 33 of their papers we have counts for

collaborators
Showing 2020Show all

14 papers · 1 filter

cs.LG2020★ 12 cited

Regret Bound Balancing and Elimination for Model Selection in Bandits and RL

Aldo Pacchiano, Christoph Dann, Claudio Gentile +1

We propose a simple model selection approach for algorithms in stochastic bandit and reinforcement learning problems. As opposed to prior work that (implicitly) assumes knowledge o…

cs.LG2020★ 3 cited

Online Model Selection for Reinforcement Learning with Function Approximation

Jonathan N. Lee, Aldo Pacchiano, Vidya Muthukumar +2

Deep reinforcement learning has achieved impressive successes yet often requires a very large amount of interaction data. This result is perhaps unsurprising, as using complicated…

cs.LG2020

Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian

Jack Parker-Holder, Luke Metz, Cinjon Resnick +6

Over the last decade, a single algorithm has changed many facets of our lives - Stochastic Gradient Descent (SGD). In the era of ever decreasing loss functions, SGD and its various…

cs.LG2020

Accelerated Message Passing for Entropy-Regularized MAP Inference

Jonathan N. Lee, Aldo Pacchiano, Peter Bartlett +1

Maximum a posteriori (MAP) inference in discrete-valued Markov random fields is a fundamental problem in machine learning that involves identifying the most likely configuration of…

cs.LG2020★ 7 cited

Stochastic Bandits with Linear Constraints

Aldo Pacchiano, Mohammad Ghavamzadeh, Peter Bartlett +1

We study a constrained contextual linear bandit setting, where the goal of the agent is to produce a sequence of policies, whose expected cumulative reward over the course of r…

cs.LG2020★ 9 cited

Regret Balancing for Bandit and RL Model Selection

Yasin Abbasi-Yadkori, Aldo Pacchiano, My Phan

We consider model selection in stochastic bandit and reinforcement learning problems. Given a set of base learning algorithms, an effective model selection strategy adapts to the b…