Practical Contextual Bandits with Regression Oracles
arXiv:1803.01088
Abstract
A major challenge in contextual bandits is to design general-purpose algorithms that are both practically useful and theoretically well-founded. We present a new technique that has the empirical and computational advantages of realizability-based approaches combined with the flexibility of agnostic methods. Our algorithms leverage the availability of a regression oracle for the value-function class, a more realistic and reasonable oracle than the classification oracles over policies typically assumed by agnostic methods. Our approach generalizes both UCB and LinUCB to far more expressive possible model classes and achieves low regret under certain distributional assumptions. In an extensive empirical evaluation, compared to both realizability-based and agnostic baselines, we find that our approach typically gives comparable or superior results.
References in corpus (3)
Cited by in corpus (18)
- A Contextual Bandit Bake-off
- Model selection for contextual bandits
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles
- Adapting multi-armed bandits policies to contextual bandits scenarios
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Federated Residual Learning
- On component interactions in two-stage recommender systems
- Upper Counterfactual Confidence Bounds: a New Optimism Principle for Contextual Bandits
- Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature
- Crush Optimism with Pessimism: Structured Bandits Beyond Asymptotic Optimality
- Adapting to Misspecification in Contextual Bandits with Offline Regression Oracles
- Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
- Tractable contextual bandits beyond realizability
- Online Sub-Sampling for Reinforcement Learning with General Function Approximation
- Rarely-switching linear bandits: optimization of causal effects for the real world
- Representation of Reinforcement Learning Policies in Reproducing Kernel Hilbert Spaces
- Boosting Image Recognition with Non-differentiable Constraints
- Going Beyond Linear RL: Sample Efficient Neural Function Approximation