Beyond Ads: Sequential Decision-Making Algorithms in Law and Public Policy
arXiv:2112.06833 · doi:10.1145/3511265.3550439
Abstract
We explore the promises and challenges of employing sequential decision-making algorithms -- such as bandits, reinforcement learning, and active learning -- in law and public policy. While such algorithms have well-characterized performance in the private sector (e.g., online advertising), the tendency to naively apply algorithms motivated by one domain, often online advertisements, can be called the "advertisement fallacy." Our main thesis is that law and public policy pose distinct methodological challenges that the machine learning community has not yet addressed. Machine learning will need to address these methodological problems to move "beyond ads." Public law, for instance, can pose multiple objectives, necessitate batched and delayed feedback, and require systems to learn rational, causal decision-making policies, each of which presents novel questions at the research frontier. We discuss a wide range of potential applications of sequential decision-making algorithms in regulation and governance, including public health, environmental protection, tax administration, occupational safety, and benefits adjudication. We use these examples to highlight research needed to render sequential decision making policy-compliant, adaptable, and effective in the public sector. We also note the potential risks of such deployments and describe how sequential decision systems can also facilitate the discovery of harms. We hope our work inspires more investigation of sequential decision making in law and public policy, which provide unique challenges for machine learning researchers with potential for significant social benefit.
Version 1 presented at Causal Inference Challenges in Sequential Decision Making: Bridging Theory and Practice (2021), a NeurIPS 2021 Workshop; Version 2 presented at the 2nd ACM Symposium on Computer Science and Law (2022) (DOI: https://dl.acm.org/doi/10.1145/3511265.3550439)
References in corpus (15)
- Unsupervised Domain Adaptation by Backpropagation
- Degenerate Feedback Loops in Recommender Systems
- Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions
- Deconvolving Feedback Loops in Recommender Systems
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Active Learning from Imperfect Labelers
- Why Adaptively Collected Data Have Negative Bias and How to Correct for It
- Active Exploration in Markov Decision Processes
- Empirical Evaluations of Active Learning Strategies in Legal Document Review
- Enhancing Environmental Enforcement with Near Real-Time Monitoring: Likelihood-Based Detection of Structural Expansion of Intensive Livestock Farms
- Estimating and Penalizing Induced Preference Shifts in Recommender Systems
- Trading off rewards and errors in multi-armed bandits
- Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature
- Design of Experiments for Stochastic Contextual Linear Bandits
- Neural Pseudo-Label Optimism for the Bank Loan Problem