Making Contextual Decisions with Low Technical Debt
arXiv:1606.03966
Abstract
Applications and systems are constantly faced with decisions that require picking from a set of actions based on contextual information. Reinforcement-based learning algorithms such as contextual bandits can be very effective in these settings, but applying them in practice is fraught with technical debt, and no general system exists that supports them completely. We address this and create the first general system for contextual learning, called the Decision Service. Existing systems often suffer from technical debt that arises from issues like incorrect data collection and weak debuggability, issues we systematically address through our ML methodology and system abstractions. The Decision Service enables all aspects of contextual bandit learning using four system abstractions which connect together in a loop: explore (the decision space), log, learn, and deploy. Notably, our new explore and log abstractions ensure the system produces correct, unbiased data, which our learner uses for online learning and to enable real-time safeguards, all in a fully reproducible manner. The Decision Service has a simple user interface and works with a variety of applications: we present two live production deployments for content recommendation that achieved click-through improvements of 25-30%, another with 18% revenue lift in the landing page, and ongoing applications in tech support and machine failure handling. The service makes real-time decisions and learns continuously and scalably, while significantly lowering technical debt.
References in corpus (6)
- Doubly Robust Policy Evaluation and Learning
- Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
- Agnostic Active Learning Without Constraints
- The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox
- Contextual Bandit Learning with Predictable Rewards
- ICE: Enabling Non-Experts to Build Models Interactively for Large-Scale Lopsided Problems
Cited by in corpus (29)
- Ray: A Distributed Framework for Emerging AI Applications
- Horizon: Facebook's Open Source Applied Reinforcement Learning Platform
- Clipper: A Low-Latency Online Prediction Serving System
- Characterizing Technical Debt and Antipatterns in AI-Based Systems: A Systematic Mapping Study
- An empirical investigation of the challenges of real-world reinforcement learning
- Reinforcement Learning Applications
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Deploying a Steered Query Optimizer in Production at Microsoft
- Federated Residual Learning
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding
- Thompson Sampling for Dynamic Pricing
- OSOM: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
- contextual: Evaluating Contextual Multi-Armed Bandit Problems in R
- Efficient First-Order Contextual Bandits: Prediction, Allocation, and Triangular Discrimination
- Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning
- Efficient Contextual Bandits with Continuous Actions
- Machine Learning Systems: A Survey from a Data-Oriented Perspective
- Lessons from Contextual Bandit Learning in a Customer Support Bot
- Empirical Likelihood for Contextual Bandits
- Data Poisoning Attacks in Contextual Bandits
- Online Evaluation of Audiences for Targeted Advertising via Bandit Experiments
- Balanced Linear Contextual Bandits
- Opportunities for Adaptive Experiments to Enable Continuous Improvement in Computer Science Education
- Tractable contextual bandits beyond realizability
- The Perils of Exploration under Competition: A Computational Modeling Approach
- Learning Accurate Decision Trees with Bandit Feedback via Quantized Gradient Descent
- Quantifying Infra-Marginality and Its Trade-off with Group Fairness
- Programming by Rewards