A Contextual-Bandit Approach to Personalized News Article Recommendation
arXiv:1003.0146 · doi:10.1145/1772690.1772758
Abstract
Personalized web services strive to adapt their services (advertisements, news articles, etc) to individual users by making use of both content and user information. Despite a few recent advances, this problem remains challenging for at least two reasons. First, web service is featured with dynamically changing pools of content, rendering traditional collaborative filtering methods inapplicable. Second, the scale of most web services of practical interest calls for solutions that are both fast in learning and computation. In this work, we model personalized recommendation of news articles as a contextual bandit problem, a principled approach in which a learning algorithm sequentially selects articles to serve users based on contextual information about the users and articles, while simultaneously adapting its article-selection strategy based on user-click feedback to maximize total user clicks. The contributions of this work are three-fold. First, we propose a new, general contextual bandit algorithm that is computationally efficient and well motivated from learning theory. Second, we argue that any bandit algorithm can be reliably evaluated offline using previously recorded random traffic. Finally, using this offline evaluation method, we successfully applied our new algorithm to a Yahoo! Front Page Today Module dataset containing over 33 million events. Results showed a 12.5% click lift compared to a standard context-free bandit algorithm, and the advantage becomes even greater when data gets more scarce.
10 pages, 5 figures
Cited by in corpus (519)
- Theory-guided Data Science: A New Paradigm for Scientific Discovery from Data
- Deep Learning based Recommender System: A Survey and New Perspectives
- Weight Uncertainty in Neural Networks
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Conservative Q-Learning for Offline Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms
- Advances and Challenges in Conversational Recommender Systems: A Survey
- How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
- Contextual Bandits with Similarity Information
- An Efficiency-boosting Client Selection Scheme for Federated Learning with Fairness Guarantee
- Towards Long-term Fairness in Recommendation
- Contextual Bandit Algorithms with Supervised Learning Guarantees
- Real-Time Bidding with Multi-Agent Reinforcement Learning in Display Advertising
- Survey on reinforcement learning for language processing
- Online Learning: A Comprehensive Survey
- Personalized News Recommendation with Context Trees
- Doubly Robust Policy Evaluation and Optimization
- Counterfactual Risk Minimization: Learning from Logged Bandit Feedback
- KuaiRand: An Unbiased Sequential Recommendation Dataset with Randomly Exposed Videos
- Neural Interactive Collaborative Filtering
- Finite-Time Analysis of Kernelised Contextual Bandits
- Reinforcement and Imitation Learning via Interactive No-Regret Learning
- Reinforcement Learning for Ridesharing: An Extended Survey
- An Efficient Bandit Algorithm for Realtime Multivariate Optimization
- MetaKG: Meta-learning on Knowledge Graph for Cold-start Recommendation
- Contextual Hybrid Session-based News Recommendation with Recurrent Neural Networks
- Learning from Logged Implicit Exploration Data
- Machine Learning Testing: Survey, Landscapes and Horizons
- Multi-Armed Bandits for Intelligent Tutoring Systems
- Science Concierge: A fast content-based recommendation system for scientific publications
- Seamlessly Unifying Attributes and Items: Conversational Recommendation for Cold-Start Users
- Distributed Clustering of Linear Bandits in Peer to Peer Networks
- Online Clustering of Bandits
- The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox
- A Survey on Contextual Multi-armed Bandits
- Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling
- Deep reinforcement learning for search, recommendation, and online advertising: a survey
- Best-Arm Identification in Linear Bandits
- Making Contextual Decisions with Low Technical Debt
- Fundamental Limits of Online and Distributed Algorithms for Statistical Learning and Estimation
- Online Interactive Collaborative Filtering Using Multi-Armed Bandit with Dependent Arms
- Learning Contextual Bandits in a Non-stationary Environment
- Leveraging Side Observations in Stochastic Bandits
- Algorithms with Logarithmic or Sublinear Regret for Constrained Contextual Bandits
- Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits
- From Bandits to Experts: On the Value of Side-Observations
- Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
- Distributed Online Learning in Social Recommender Systems
- What Doubling Tricks Can and Can't Do for Multi-Armed Bandits
- Spectral bandits for smooth graph functions
- RecSim: A Configurable Simulation Platform for Recommender Systems
- A Contextual Bandit Bake-off
- Learning Graph Meta Embeddings for Cold-Start Ads in Click-Through Rate Prediction
- Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions
- An Actor-Critic Contextual Bandit Algorithm for Personalized Mobile Health Interventions
- Reinforcement Learning Applications
- Statistical Decision Making for Optimal Budget Allocation in Crowd Labeling
- Differentially-Private Federated Linear Bandits
- Knowledge-guided Deep Reinforcement Learning for Interactive Recommendation
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
- Comparison-based Conversational Recommender System with Relative Bandit Feedback
- Off-policy evaluation for slate recommendation
- Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning
- Context-Aware Hierarchical Online Learning for Performance Maximization in Mobile Crowdsourcing
- Multi-objective Contextual Multi-armed Bandit with a Dominant Objective
- Carousel Personalization in Music Streaming Apps with Contextual Bandits
- Offline Contextual Multi-armed Bandits for Mobile Health Interventions: A Case Study on Emotion Regulation
- Hedging the Drift: Learning to Optimize under Non-Stationarity
- Neural Thompson Sampling
- Thompson Sampling for Budgeted Multi-armed Bandits
- Beyond Preferences in AI Alignment
- Sliding Spectrum Decomposition for Diversified Recommendation
- A Practical Algorithm for Multiplayer Bandits when Arm Means Vary Among Players
- Sequential Batch Learning in Finite-Action Linear Contextual Bandits
- Scalable Generalized Linear Bandits: Online Computation and Hashing
- Statistical Inference for Online Decision-Making: In a Contextual Bandit Setting
- Cascading Hybrid Bandits: Online Learning to Rank for Relevance and Diversity
- Multi-Task Learning for Contextual Bandits
- The Digital Transformation in Health: How AI Can Improve the Performance of Health Systems
- Visual Intelligence through Human Interaction
- Modeling Rabbit-Holes on YouTube
- Sketch-based Creativity Support Tools using Deep Learning
- Online Learning under Delayed Feedback
- Exploring Personalized Health Support through Data-Driven, Theory-Guided LLMs: A Case Study in Sleep Health
- Estimation Considerations in Contextual Bandits
- On Multi-Armed Bandit Designs for Dose-Finding Clinical Trials
- Online Machine Learning in Big Data Streams
- From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
- Adapting multi-armed bandits policies to contextual bandits scenarios
- Reinforcement Learning in Modern Biostatistics: Constructing Optimal Adaptive Interventions
- Exploration in Interactive Personalized Music Recommendation: A Reinforcement Learning Approach
- Meta Dynamic Pricing: Transfer Learning Across Experiments
- Bootstrapping Upper Confidence Bound
- From Predictions to Prescriptions in Multistage Optimization Problems
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical Activity
- Agnostic Q-learning with Function Approximation in Deterministic Systems: Tight Bounds on Approximation Error and Sample Complexity
- Nearly Minimax Optimal Reinforcement Learning for Linear Mixture Markov Decision Processes
- Doubly-Robust Lasso Bandit
- Model Selection for Offline Reinforcement Learning: Practical Considerations for Healthcare Settings
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- RELEAF: An Algorithm for Learning and Exploiting Relevance
- Benchmarks for Deep Off-Policy Evaluation
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Federated Bandit: A Gossiping Approach
- Beyond Ads: Sequential Decision-Making Algorithms in Law and Public Policy
- Policy Gradients for Contextual Recommendations
- Logarithmic Regret for Reinforcement Learning with Linear Function Approximation
- Streaming kernel regression with provably adaptive mean, variance, and regularization
- User Tampering in Reinforcement Learning Recommender Systems
- A Practical Method for Solving Contextual Bandit Problems Using Decision Trees
- Action-Manipulation Attacks Against Stochastic Bandits: Attacks and Defense
- Neural Contextual Bandits with Deep Representation and Shallow Exploration
- Optimal Best-arm Identification in Linear Bandits
- Counterfactual Estimation and Optimization of Click Metrics for Search Engines
- Adversarial Attacks on Linear Contextual Bandits
- Fully adaptive algorithm for pure exploration in linear bandits
- Statistical Inference for Online Decision Making via Stochastic Gradient Descent
- Inference for Batched Bandits
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization
- An Efficient Pessimistic-Optimistic Algorithm for Stochastic Linear Bandits with General Constraints
- Efficient Change-Point Detection for Tackling Piecewise-Stationary Bandits
- Almost Optimal Algorithms for Linear Stochastic Bandits with Heavy-Tailed Payoffs
- Stochastic Bandits with Context Distributions
- Improving drug sensitivity predictions in precision medicine through active expert knowledge elicitation
- Partially Observable Markov Decision Process for Recommender Systems
- Multi-Agent Multi-Armed Bandits with Limited Communication
- Federated Linear Contextual Bandits
- Optimal No-regret Learning in Repeated First-price Auctions
- Instructions and Guide for Diagnostic Questions: The NeurIPS 2020 Education Challenge
- Stochastic Contextual Bandits with Known Reward Functions
- Federated Multi-armed Bandits with Personalization
- Learning Triggers for Heterogeneous Treatment Effects
- Improving offline evaluation of contextual bandit algorithms via bootstrapping techniques
- Hierarchical Exploration for Accelerating Contextual Bandits
- Impatient Bandits: Optimizing Recommendations for the Long-Term Without Delay
- BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback
- Simple Regret Minimization for Contextual Bandits
- SHAPFUZZ: Efficient Fuzzing via Shapley-Guided Byte Selection
- Variational inference for the multi-armed contextual bandit
- Locally Differentially Private (Contextual) Bandits Learning
- Output-Weighted Sampling for Multi-Armed Bandits with Extreme Payoffs
- Post-Contextual-Bandit Inference
- Bilinear Bandits with Low-rank Structure
- Automated Creative Optimization for E-Commerce Advertising
- Off-Policy Evaluation and Learning for External Validity under a Covariate Shift
- Graphical Models for Bandit Problems
- Reinforcement Learning for Personalized Dialogue Management
- We Know What You Want: An Advertising Strategy Recommender System for Online Advertising
- Stochastic Structured Prediction under Bandit Feedback
- Garbage In, Reward Out: Bootstrapping Exploration in Multi-Armed Bandits
- Graph Neural News Recommendation with Long-term and Short-term Interest Modeling
- Semiparametric Contextual Bandits
- Ranked bandits in metric spaces: learning optimally diverse rankings over large document collections
- Regret Bound Balancing and Elimination for Model Selection in Bandits and RL
- Greybox fuzzing as a contextual bandits problem
- A Bandit Approach to Posterior Dialog Orchestration Under a Budget
- A Context-aware Radio Resource Management in Heterogeneous Virtual RANs
- New Insights into Bootstrapping for Bandits
- Exploration in Online Advertising Systems with Deep Uncertainty-Aware Learning
- Speaker Diarization as a Fully Online Learning Problem in MiniVox
- Simulating Non Stationary Operators in Search Algorithms
- Graph Clustering Bandits for Recommendation
- Latent Bandits Revisited
- Bandit Online Learning with Unknown Delays
- Do Offline Metrics Predict Online Performance in Recommender Systems?
- State Encoders in Reinforcement Learning for Recommendation: A Reproducibility Study
- Selfish Robustness and Equilibria in Multi-Player Bandits
- On component interactions in two-stage recommender systems
- Off-Policy Risk Assessment in Contextual Bandits
- Graph Neural Bandits
- Deep Contextual Multi-armed Bandits
- Online learning with Corrupted context: Corrupted Contextual Bandits
- Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental Health
- A Neural Networks Committee for the Contextual Bandit Problem
- Deep Neural Linear Bandits: Overcoming Catastrophic Forgetting through Likelihood Matching
- PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits
- Meta-Learning for Contextual Bandit Exploration
- Large-scale Interactive Recommendation with Tree-structured Policy Gradient
- EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits
- On Ensuring that Intelligent Machines Are Well-Behaved
- A Re-classification of Information Seeking Tasks and Their Computational Solutions
- Fairness of Exposure in Stochastic Bandits
- Optimal Baseline Corrections for Off-Policy Contextual Bandits
- RecInDial: A Unified Framework for Conversational Recommendation with Pretrained Language Models
- Sequential Experimental Design for Transductive Linear Bandits
- Beyond Personalization: Social Content Recommendation for Creator Equality and Consumer Satisfaction
- Towards Practical Lipschitz Bandits
- Reinforcement Learning for Combining Search Methods in the Calibration of Economic ABMs
- Online learning in MDPs with side information
- Online Learning in Contextual Bandits using Gated Linear Networks
- Cold-start Problems in Recommendation Systems via Contextual-bandit Algorithms
- An Efficient Algorithm For Generalized Linear Bandit: Online Stochastic Gradient Descent and Thompson Sampling
- The closed loop between opinion formation and personalised recommendations
- Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks
- Deep density networks and uncertainty in recommender systems
- Hierarchical Adaptive Contextual Bandits for Resource Constraint based Recommendation
- contextual: Evaluating Contextual Multi-Armed Bandit Problems in R
- Designing and Deploying Online Field Experiments
- OSOM: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
- Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication
- Real-Time Optimization Of Web Publisher RTB Revenues
- Recommender System for News Articles using Supervised Learning
- Stochastic Linear Bandits Robust to Adversarial Attacks
- Kernel Methods for Cooperative Multi-Agent Contextual Bandits
- Machine Translation System Selection from Bandit Feedback
- Debiased Off-Policy Evaluation for Recommendation Systems
- Gaussian Process Optimization with Adaptive Sketching: Scalable and No Regret
- Getting too personal(ized): The importance of feature choice in online adaptive algorithms
- Corralling a Band of Bandit Algorithms
- Stochastic Bandits with Linear Constraints
- Online Semi-Supervised Learning with Bandit Feedback
- Empirical Bayes Regret Minimization
- Beyond the Click-Through Rate: Web Link Selection with Multi-level Feedback
- Diffusion Approximations for Thompson Sampling in the Small Gap Regime
- Conversational Dueling Bandits in Generalized Linear Models
- Local Differential Privacy for Bayesian Optimization
- Bandit Algorithms for Precision Medicine
- Hidden Incentives for Auto-Induced Distributional Shift
- Unknowable Manipulators: Social Network Curator Algorithms
- Markov Decision Processes with Continuous Side Information
- Federated Multi-Armed Bandits
- Structured Linear Contextual Bandits: A Sharp and Geometric Smoothed Analysis
- A Smoothed Analysis of the Greedy Algorithm for the Linear Contextual Bandit Problem
- Best-Arm Identification in Correlated Multi-Armed Bandits
- Sequential Monte Carlo Bandits
- Improved Algorithm on Online Clustering of Bandits
- Charging control of electric vehicles using contextual bandits considering the electrical distribution grid
- RL4health: Crowdsourcing Reinforcement Learning for Knee Replacement Pathway Optimization
- Ed-Fed: A generic federated learning framework with resource-aware client selection for edge devices
- Managing Risk of Bidding in Display Advertising
- Efficient Contextual Bandits with Continuous Actions
- Human Language Modeling
- Adaptive Exploration in Linear Contextual Bandit
- Meta-Learning Bandit Policies by Gradient Ascent
- Chameleon: Foundation Models for Fairness-aware Multi-modal Data Augmentation to Enhance Coverage of Minorities
- Accurate Inference for Adaptive Linear Models
- On the Power of Multitask Representation Learning in Linear MDP
- The Intrinsic Robustness of Stochastic Bandits to Strategic Manipulation
- Empirical Likelihood for Contextual Bandits
- Nonstationary Stochastic Multiarmed Bandits: UCB Policies and Minimax Regret
- Nonparametric Stochastic Contextual Bandits
- Hyper-parameter Tuning for the Contextual Bandit
- Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation
- Fatigue-Aware Ad Creative Selection
- Group Fairness in Bandit Arm Selection
- Towards Validating Long-Term User Feedbacks in Interactive Recommendation Systems
- Reward Constrained Interactive Recommendation with Natural Language Feedback
- Handling Cold-Start Collaborative Filtering with Reinforcement Learning
- Mobile Recommender Systems Methods: An Overview
- Meta-learning with Stochastic Linear Bandits
- Multi-Objective Generalized Linear Bandits
- Contextual Bandit with Missing Rewards
- Sparsity-Agnostic Lasso Bandit
- Lessons from Contextual Bandit Learning in a Customer Support Bot
- Structured Stochastic Linear Bandits
- Thompson Sampling for Contextual Bandit Problems with Auxiliary Safety Constraints
- Contextual Bandits for adapting to changing User preferences over time
- Dynamic Learning of Sequential Choice Bandit Problem under Marketing Fatigue
- Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio Estimation
- Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits
- Thresholded Lasso Bandit
- Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
- Uplift Modeling for Multiple Treatments with Cost Optimization
- Social Explorative Attention based Recommendation for Content Distribution Platforms
- Mitigating Bias in Adaptive Data Gathering via Differential Privacy
- Knowledge Infused Policy Gradients for Adaptive Pandemic Control
- Deep Reinforcement Learning for Adaptive Learning Systems
- MergeDTS: A Method for Effective Large-Scale Online Ranker Evaluation
- A Linear Bandit for Seasonal Environments
- InferLine: ML Prediction Pipeline Provisioning and Management for Tight Latency Objectives
- Randomized Allocation with Nonparametric Estimation for Contextual Multi-Armed Bandits with Delayed Rewards
- Adaptive Experimentation with Delayed Binary Feedback
- Adaptive Keywords Extraction with Contextual Bandits for Advertising on Parked Domains
- Exploration in two-stage recommender systems
- Nearly Dimension-Independent Sparse Linear Bandit over Small Action Spaces via Best Subset Selection
- Dynamic Product Image Generation and Recommendation at Scale for Personalized E-commerce
- Unified Dual-Intent Translation for Joint Modeling of Search and Recommendation
- Rapidly Personalizing Mobile Health Treatment Policies with Limited Data
- Asynchronous Parallel Empirical Variance Guided Algorithms for the Thresholding Bandit Problem
- Modeling User Exposure in Recommendation
- Learning Optimal Data Augmentation Policies via Bayesian Optimization for Image Classification Tasks
- Risk-Aware Continuous Control with Neural Contextual Bandits
- Multi-Armed Bandits with Fairness Constraints for Distributing Resources to Human Teammates
- A sequential transit network design algorithm with optimal learning under correlated beliefs
- Quantum contextual bandits and recommender systems for quantum data
- Autonomous Drug Design with Multi-Armed Bandits
- Neural Bandit with Arm Group Graph
- Nonparametric Contextual Bandits in an Unknown Metric Space
- ASAC: Active Sensing using Actor-Critic models
- Adapting to Misspecification in Contextual Bandits with Offline Regression Oracles
- Efficient and Robust Algorithms for Adversarial Linear Contextual Bandits
- Impact of Representation Learning in Linear Bandits
- Robust Stochastic Linear Contextual Bandits Under Adversarial Attacks
- A Practical Guide of Off-Policy Evaluation for Bandit Problems
- Data Poisoning Attacks in Contextual Bandits
- Refining Recency Search Results with User Click Feedback
- Sample-efficient Nonstationary Policy Evaluation for Contextual Bandits
- Evaluation of Explore-Exploit Policies in Multi-result Ranking Systems
- Discovering Valuable Items from Massive Data
- Online Learning to Estimate Warfarin Dose with Contextual Linear Bandits
- Reducing Exploration of Dying Arms in Mortal Bandits
- When Collaborative Filtering Meets Reinforcement Learning
- PAC Identification of Many Good Arms in Stochastic Multi-Armed Bandits
- Thompson Sampling for a Fatigue-aware Online Recommendation System
- Balanced Linear Contextual Bandits
- Stochastic Linear Contextual Bandits with Diverse Contexts
- Online Preselection with Context Information under the Plackett-Luce Model
- Sublinear Optimal Policy Value Estimation in Contextual Bandits
- An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing Platforms
- Deep Bayesian Bandits: Exploring in Online Personalized Recommendations
- Linear Contextual Bandits with Adversarial Corruptions
- Estimation of Warfarin Dosage with Reinforcement Learning
- X2T: Training an X-to-Text Typing Interface with Online Learning from User Feedback
- Efficient Online Bayesian Inference for Neural Bandits
- Dynamic Global Sensitivity for Differentially Private Contextual Bandits
- A Field Test of Bandit Algorithms for Recommendations: Understanding the Validity of Assumptions on Human Preferences in Multi-armed Bandits
- Personalized Product Assortment with Real-time 3D Perception and Bayesian Payoff Estimation
- Max-Utility Based Arm Selection Strategy For Sequential Query Recommendations
- Finite-time Analysis of Globally Nonstationary Multi-Armed Bandits
- Episodic Multi-armed Bandits
- Tractable contextual bandits beyond realizability
- Contextual Bandits with Side-Observations
- The Gossiping Insert-Eliminate Algorithm for Multi-Agent Bandits
- Edge Computing in the Dark: Leveraging Contextual-Combinatorial Bandit and Coded Computing
- Old Dog Learns New Tricks: Randomized UCB for Bandit Problems
- Online Evaluation of Audiences for Targeted Advertising via Bandit Experiments
- ADARES: Adaptive Resource Management for Virtual Machines
- Explore-Exploit: A Framework for Interactive and Online Learning
- Contextual Bandits with Stochastic Experts
- Collaborative Filtering Bandits
- Recommandation mobile, sensible au contexte de contenus évolutifs: Contextuel-E-Greedy
- Computational Causal Inference
- Thompson Sampling for Noncompliant Bandits
- Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity
- Decision Making Problems with Funnel Structure: A Multi-Task Learning Approach with Application to Email Marketing Campaigns
- Context-Aware Bandits
- Dimension Reduction in Contextual Online Learning via Nonparametric Variable Selection
- Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits
- Online Algorithm for Unsupervised Sequential Selection with Contextual Information
- Off-Policy Evaluation of Bandit Algorithm from Dependent Samples under Batch Update Policy
- Stochastic Lipschitz Q-Learning
- Online Batch Decision-Making with High-Dimensional Covariates
- R-UCB: a Contextual Bandit Algorithm for Risk-Aware Recommender Systems
- Policy Optimization with Model-based Explorations
- Exploitation Over Exploration: Unmasking the Bias in Linear Bandit Recommender Offline Evaluation
- Online Limited Memory Neural-Linear Bandits with Likelihood Matching
- Active Collaborative Sensing for Energy Breakdown
- A Hidden Markov Restless Multi-armed Bandit Model for Playout Recommendation Systems
- Robust Actor-Critic Contextual Bandit for Mobile Health (mHealth) Interventions
- Weighted Gaussian Process Bandits for Non-stationary Environments
- Contextual Bandit Applications in Customer Support Bot
- Improved Confidence Bounds for the Linear Logistic Model and Applications to Linear Bandits
- Incentivizing Exploration in Linear Bandits under Information Gap
- Multi-Faceted Ranking of News Articles using Post-Read Actions
- Interactive Exploration and Discovery of Scientific Publications with PubVis
- Federated Multi-Armed Bandits Under Byzantine Attacks
- Contextual Bandit with Adaptive Feature Extraction
- Reinforcement Learning for Strategic Recommendations
- Bandits with Knapsacks beyond the Worst-Case
- Decentralized, Adaptive, Look-Ahead Particle Filtering
- Multi-Armed Bandits for Correlated Markovian Environments with Smoothed Reward Feedback
- Towards the D-Optimal Online Experiment Design for Recommender Selection
- Exploiting Structure of Uncertainty for Efficient Matroid Semi-Bandits
- Conservative Contextual Combinatorial Cascading Bandit
- Learning to Use Learners' Advice
- Online learning in bandits with predicted context
- Automatic Representation for Lifetime Value Recommender Systems
- Task Offloading and Replication for Vehicular Cloud Computing: A Multi-Armed Bandit Approach
- Budget-constrained Edge Service Provisioning with Demand Estimation via Bandit Learning
- Interactive Visualization Recommendation with Hier-SUCB
- Action Centered Contextual Bandits
- Predicting Counterfactuals from Large Historical Data and Small Randomized Trials
- Freshness-Aware Thompson Sampling
- Actively Learning to Attract Followers on Twitter
- Multitask Bandit Learning Through Heterogeneous Feedback Aggregation
- Concentration bounds for temporal difference learning with linear function approximation: The case of batch data and uniform sampling
- Solving Multi-Arm Bandit Using a Few Bits of Communication
- Toward an Integrated Framework for Automated Development and Optimization of Online Advertising Campaigns
- Combining Online Learning and Offline Learning for Contextual Bandits with Deficient Support
- An Arm-Wise Randomization Approach to Combinatorial Linear Semi-Bandits
- Optimized Recommender Systems with Deep Reinforcement Learning
- AutoML for Contextual Bandits
- BanditMF: Multi-Armed Bandit Based Matrix Factorization Recommender System
- Banker Online Mirror Descent
- PHYRE: A New Benchmark for Physical Reasoning
- Thompson Sampling Algorithms for Cascading Bandits
- Efficient Inference Without Trading-off Regret in Bandits: An Allocation Probability Test for Thompson Sampling
- Self-Concordant Analysis of Generalized Linear Bandits with Forgetting
- Metric-Free Individual Fairness with Cooperative Contextual Bandits
- Improving Offline Contextual Bandits with Distributional Robustness
- Combinatorial Bandits for Incentivizing Agents with Dynamic Preferences
- Targeted Advertising on Social Networks Using Online Variational Tensor Regression
- Adversarial Linear Contextual Bandits with Graph-Structured Side Observations
- Fast Physical Activity Suggestions: Efficient Hyperparameter Learning in Mobile Health
- Scalable Multi-Class Bayesian Support Vector Machines for Structured and Unstructured Data
- Deep Reinforcement Learning for Personalized Search Story Recommendation
- Counterfactual Contextual Multi-Armed Bandit: a Real-World Application to Diagnose Apple Diseases
- Uniform-PAC Bounds for Reinforcement Learning with Linear Function Approximation
- Challenges in Statistical Analysis of Data Collected by a Bandit Algorithm: An Empirical Exploration in Applications to Adaptively Randomized Experiments
- Adaptive Doubly Robust Estimator from Non-stationary Logging Policy under a Convergence of Average Probability
- UCBoost: A Boosting Approach to Tame Complexity and Optimality for Stochastic Bandits
- Hawkes Process Multi-armed Bandits for Disaster Search and Rescue
- Inverse Contextual Bandits: Learning How Behavior Evolves over Time
- Adversarial Bandits with Multi-User Delayed Feedback: Theory and Application
- Acting Selfish for the Good of All: Contextual Bandits for Resource-Efficient Transmission of Vehicular Sensor Data
- New Classes of the Greedy-Applicable Arm Feature Distributions in the Sparse Linear Bandit Problem
- Linear Contextual Bandits with Hybrid Payoff: Revisited
- Non-Stationary Off-Policy Optimization
- Personalized Web Search
- Incentivized Exploration for Multi-Armed Bandits under Reward Drift
- Syndicated Bandits: A Framework for Auto Tuning Hyper-parameters in Contextual Bandit Algorithms
- Rarely-switching linear bandits: optimization of causal effects for the real world
- Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack
- Contextual-Bandit Based Personalized Recommendation with Time-Varying User Interests
- Three Methods for Training on Bandit Feedback
- Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual Bandits
- Evolution of the user's content: An Overview of the state of the art
- Universal and data-adaptive algorithms for model selection in linear contextual bandits
- Predicting Tomorrow's Headline using Today's Twitter Deliberations
- Causal Simulations for Uplift Modeling
- AdaLinUCB: Opportunistic Learning for Contextual Bandits
- A Load Balanced Recommendation Approach
- Recommending Dream Jobs in a Biased Real World
- Lipschitz Bandit Optimization with Improved Efficiency
- Optimal Recommendation to Users that React: Online Learning for a Class of POMDPs
- Algorithms for Linear Bandits on Polyhedral Sets
- Multi-armed Bandit Requiring Monotone Arm Sequences
- Differentially Private Online Learning for Cloud-Based Video Recommendation with Multimedia Big Data in Social Networks
- Control Variates for Slate Off-Policy Evaluation
- Predicting next shopping stage using Google Analytics data for E-commerce applications
- Contextual Online Learning for Multimedia Content Aggregation
- Bandit Learning for Diversified Interactive Recommendation
- An Active Learning Framework for Efficient Robust Policy Search
- Contextual Recommendations and Low-Regret Cutting-Plane Algorithms
- Robust Contextual Bandit via the Capped- norm
- Koopman Q-learning: Offline Reinforcement Learning via Symmetries of Dynamics
- Sequential Matrix Completion
- Online Semi-Supervised Learning in Contextual Bandits with Episodic Reward
- Offline Contextual Bandits with Overparameterized Models
- DART: aDaptive Accept RejecT for non-linear top-K subset identification
- A Simple Geometric Method for Cross-Lingual Linguistic Transformations with Pre-trained Autoencoders
- Stochastic Multi-Armed Bandits with Control Variates
- DTR Bandit: Learning to Make Response-Adaptive Decisions With Low Regret
- Contextual Bandits with Sparse Data in Web setting
- Guaranteed Fixed-Confidence Best Arm Identification in Multi-Armed Bandits: Simple Sequential Elimination Algorithms
- Active Offline Policy Selection
- Streamlined Empirical Bayes Fitting of Linear Mixed Models in Mobile Health
- Decomposition-Coordination Method for Finite Horizon Bandit Problems
- Neural Contextual Bandits without Regret
- Bandits with Stochastic Experts: Constant Regret, Empirical Experts and Episodes
- Thompson Sampling for Bandits with Clustered Arms
- Apple Tasting Revisited: Bayesian Approaches to Partially Monitored Online Binary Classification
- Deep Upper Confidence Bound Algorithm for Contextual Bandit Ranking of Information Selection
- Linear Bandit algorithms using the Bootstrap
- Privacy-Preserving Bandits
- Contributions to Representation Learning with Graph Autoencoders and Applications to Music Recommendation
- Pessimistic Off-Policy Optimization for Learning to Rank
- Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures
- Leveraging heterogeneous spillover in maximizing contextual bandit rewards
- A Time and Space Efficient Algorithm for Contextual Linear Bandits
- Active Inference in Contextual Multi-Armed Bandits for Autonomous Robotic Exploration
- Learning Orthogonal Projections in Linear Bandits
- The Nah Bandit: Modeling User Non-compliance in Recommendation Systems
- MBExplainer: Multilevel bandit-based explanations for downstream models with augmented graph embeddings
- Spectral bandits for smooth graph functions with applications in recommender systems
- AdSEE: Investigating the Impact of Image Style Editing on Advertisement Attractiveness
- Recommending with Recommendations
- An Adversarial Learning based Multi-Step Spoken Language Understanding System through Human-Computer Interaction
- The Use of Bandit Algorithms in Intelligent Interactive Recommender Systems
- A Map of Bandits for E-commerce
- GuideBoot: Guided Bootstrap for Deep Contextual Bandits
- An Empirical Analysis on Transparent Algorithmic Exploration in Recommender Systems
- Neural News Recommendation with Collaborative News Encoding and Structural User Encoding
- Generalized Translation and Scale Invariant Online Algorithm for Adversarial Multi-Armed Bandits
- Residual Overfit Method of Exploration
- Minimizing the Age of Information from Sensors with Common Observations
- Sayer: Using Implicit Feedback to Optimize System Policies
- D2RLIR : an improved and diversified ranking function in interactive recommendation systems based on deep reinforcement learning
- Spatio-temporal Edge Service Placement: A Bandit Learning Approach
- The Impact of Batch Learning in Stochastic Bandits
- Robust Counterfactual Inferences using Feature Learning and their Applications
- DORB: Dynamically Optimizing Multiple Rewards with Bandits
- Unbiased Estimation of the Value of an Optimized Policy
- Neural News Recommendation with Negative Feedback
- Client-Based Intelligence for Resource Efficient Vehicular Big Data Transfer in Future 6G Network
- Robust Bandit Learning with Imperfect Context
- An Opportunistic Bandit Approach for User Interface Experimentation
- A Smoothed Analysis of Online Lasso for the Sparse Linear Contextual Bandit Problem
- Greedy Bandits with Sampled Context
- On Adaptive Estimation for Dynamic Bernoulli Bandits
- Green Offloading in Fog-Assisted IoT Systems: An Online Perspective Integrating Learning and Control
- Unifying Clustered and Non-stationary Bandits
- Improved Algorithms for Conservative Exploration in Bandits
- AI Online Filters to Real World Image Recognition
- Understanding individual behaviour: from virtual to physical patterns
- Multiscale Non-stationary Stochastic Bandits
- Survey Bandits with Regret Guarantees
- A Real-Time Framework for Task Assignment in Hyperlocal Spatial Crowdsourcing
- Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
- The Lingering of Gradients: Theory and Applications
- An Unbiased Data Collection and Content Exploitation/Exploration Strategy for Personalization
- CBA: Contextual Quality Adaptation for Adaptive Bitrate Video Streaming (Extended Version)
- The Assistive Multi-Armed Bandit
- Contextual Bandits with Random Projection
- A Methodology for Discovering how to Adaptively Personalize to Users using Experimental Comparisons
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy
- Random Forest for the Contextual Bandit Problem - extended version
- Local Optimality of User Choices and Collaborative Competitive Filtering
- Towards Bursting Filter Bubble via Contextual Risks and Uncertainties
- Multi-level Feedback Web Links Selection Problem: Learning and Optimization
- Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays
- The Adaptive Doubly Robust Estimator for Policy Evaluation in Adaptive Experiments and a Paradox Concerning Logging Policy
- Diversity-Preserving K-Armed Bandits, Revisited
- Online Action Learning in High Dimensions: A Conservative Perspective
- Towards Fundamental Limits of Multi-armed Bandits with Random Walk Feedback