activity
20172022
most citedSpectrally-normalized margin bounds for neural networks

174 citations · 313 across the 14 of their papers we have counts for

collaborators

24 papers

cs.LG202121 cited

Adapting to Misspecification in Contextual Bandits

Dylan J. Foster, Claudio Gentile, Mehryar Mohri +1

A major research direction in contextual bandits is to develop algorithms that are computationally efficient, yet support flexible, general-purpose function approximation. Algorith…

cs.LG20218 cited

Efficient First-Order Contextual Bandits: Prediction, Allocation, and Triangular Discrimination

Dylan J. Foster, Akshay Krishnamurthy

A recurring theme in statistical learning, online learning, and beyond is that faster convergence rates are possible for problems with low noise, often quantified by the performanc…

cs.LG202122 cited

Independent Policy Gradient Methods for Competitive Reinforcement Learning

Constantinos Daskalakis, Dylan J. Foster, Noah Golowich

We obtain global, non-asymptotic convergence guarantees for independent learning algorithms in competitive reinforcement learning settings with two agents (i.e., zero-sum stochasti…

cs.LG2020

Learning the Linear Quadratic Regulator from Nonlinear Observations

Zakaria Mhammedi, Dylan J. Foster, Max Simchowitz +5

We introduce a new problem setting for continuous control called the LQR with Rich Observations, or RichLQR. In our setting, the environment is summarized by a low-dimensional cont…

cs.LG2020

Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective

Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi +1

In the classical multi-armed bandit problem, instance-dependent algorithms attain improved performance on "easy" problems with a gap between the best and second-best arm. Are simil…

cs.LG20202 cited

Tight Bounds on Minimax Regret under Logarithmic Loss via Self-Concordance

Blair Bilodeau, Dylan J. Foster, Daniel M. Roy

We consider the classical problem of sequential probability assignment under logarithmic loss while competing against an arbitrary, potentially nonparametric class of experts. We o…