Real-Time Bidding with Multi-Agent Reinforcement Learning in Display Advertising
arXiv:1802.09756 · doi:10.1145/3269206.3272021
Abstract
Real-time advertising allows advertisers to bid for each impression for a visiting user. To optimize specific goals such as maximizing revenue and return on investment (ROI) led by ad placements, advertisers not only need to estimate the relevance between the ads and user's interests, but most importantly require a strategic response with respect to other advertisers bidding in the market. In this paper, we formulate bidding optimization with multi-agent reinforcement learning. To deal with a large number of advertisers, we propose a clustering method and assign each cluster with a strategic bidding agent. A practical Distributed Coordinated Multi-Agent Bidding (DCMAB) has been proposed and implemented to balance the tradeoff between the competition and cooperation among advertisers. The empirical study on our industry-scaled real-world data has demonstrated the effectiveness of our methods. Our results show cluster-based bidding would largely outperform single-agent and bandit approaches, and the coordinated bidding achieves better overall objectives than purely self-interested bidding agents.
References in corpus (9)
- Continuous control with deep reinforcement learning
- A Contextual-Bandit Approach to Personalized News Article Recommendation
- Learning to Communicate with Deep Multi-Agent Reinforcement Learning
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- Multi-agent Reinforcement Learning in Sequential Social Dilemmas
- Real-Time Bidding by Reinforcement Learning in Display Advertising
- Optimized Cost per Click in Taobao Display Advertising
- Smart Pacing for Effective Online Ad Campaign Optimization
- LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions
Cited by in corpus (32)
- Learning Tree-based Deep Model for Recommender Systems
- Deep reinforcement learning for search, recommendation, and online advertising: a survey
- Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling
- Jointly Learning to Recommend and Advertise
- Joint Optimization of Tree-based Index and Deep Model for Recommender Systems
- A Cooperative-Competitive Multi-Agent Framework for Auto-bidding in Online Advertising
- Statistical Inference of the Value Function for Reinforcement Learning in Infinite Horizon Settings
- Dynamic Programming Principles for Mean-Field Controls with Learning
- Deep Reinforcement Learning with Robust and Smooth Policy
- Large-scale Interactive Recommendation with Tree-structured Policy Gradient
- Hierarchically Constrained Adaptive Ad Exposure in Feeds
- The PlayStation Reinforcement Learning Environment (PSXLE)
- Learning Mean-Field Games
- Flatland Competition 2020: MAPF and MARL for Efficient Train Coordination on a Grid World
- MoTiAC: Multi-Objective Actor-Critics for Real-Time Bidding
- Distributed TD(0) with Almost No Communication
- Risk-Aware Bid Optimization for Online Display Advertisement
- Interpretable Imitation Learning with Dynamic Causal Relations
- Mean-Field Multi-Agent Reinforcement Learning: A Decentralized Network Approach
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Compression and Localization in Reinforcement Learning for ATARI Games
- Causal Multi-Agent Reinforcement Learning: Review and Open Problems
- Accelerating Offline Reinforcement Learning Application in Real-Time Bidding and Recommendation: Potential Use of Simulation
- Temporal Difference Learning as Gradient Splitting
- Learning Adaptive Display Exposure for Real-Time Advertising
- Policy Gradient Methods Find the Nash Equilibrium in N-player General-sum Linear-quadratic Games
- Expert-Guided Diffusion Planner for Auto-Bidding
- Infer Your Enemies and Know Yourself, Learning in Real-Time Bidding with Partially Observable Opponents
- Multi-Agent Cooperative Bidding Games for Multi-Objective Optimization in e-Commercial Sponsored Search
- Truncation-Free Matching System for Display Advertising at Alibaba
- Learning to Advertise for Organic Traffic Maximization in E-Commerce Product Feeds
- Improving On-policy Learning with Statistical Reward Accumulation