Publications (12)
A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design
Rui Ai, Boxiang Lyu, Zhaoran Wang +2
We study reserve price optimization in multi-phase second price auctions, where the seller's prior actions affect the bidders' later valuations through a Markov Decision Process (M…
Pairwise Ranking Losses of Click-Through Rates Prediction for Welfare Maximization in Ad Auctions
Boxiang Lyu, Zhe Feng, Zachary Robertson +1
We study the design of loss functions for click-through rates (CTR) to optimize (social) welfare in advertising auctions. Existing works either only focus on CTR predictions withou…
Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting
Zhongjian Qiao, Jiafei Lyu, Boxiang Lyu +3
Model-based offline reinforcement learning (RL) aims to enhance offline RL with a dynamics model that facilitates policy exploration. However, \textit{model exploitation} could occ…
Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach
Shuang Qiu, Boxiang Lyu, Qinglin Meng +3
Dynamic mechanism design studies how mechanism designers should allocate resources among agents in a time-varying environment. We consider the problem where the agents interact wit…
An Instrumental Value for Data Production and its Application to Data Pricing
Rui Ai, Boxiang Lyu, Zhaoran Wang +2
How much value does a dataset or a data production process have to an agent who wishes to use the data to assist decision-making? This is a fundamental question towards understandi…
Pessimism meets VCG: Learning Dynamic Mechanism Design via Offline Reinforcement Learning
Boxiang Lyu, Zhaoran Wang, Mladen Kolar +1
Dynamic mechanism design has garnered significant attention from both computer scientists and economists in recent years. By allowing agents to interact with the seller over multip…
Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
Shuang Qiu, Dake Zhang, Rui Yang +2
This paper investigates multi-objective reinforcement learning (MORL), which focuses on learning Pareto optimal policies in the presence of multiple reward functions. Despite MORL'…
L-SVRG and L-Katyusha with Adaptive Sampling
Boxin Zhao, Boxiang Lyu, Mladen Kolar
Stochastic gradient-based optimization methods, such as L-SVRG and its accelerated variant L-Katyusha (Kovalev et al., 2020), are widely used to train machine learning models.The t…
Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling Algorithm
Boxin Zhao, Boxiang Lyu, Raul Castro Fernandez +1
High-quality machine learning models are dependent on access to high-quality training data. When the data are not already available, it is tedious and costly to obtain them. Data m…
One Policy is Enough: Parallel Exploration with a Single Policy is Near-Optimal for Reward-Free Reinforcement Learning
Pedro Cisneros-Velarde, Boxiang Lyu, Sanmi Koyejo +1
Although parallelism has been extensively used in reinforcement learning (RL), the quantitative effects of parallel exploration are not well understood theoretically. We study the…
Personalized Federated Learning with Multiple Known Clusters
Boxiang Lyu, Filip Hanzely, Mladen Kolar
We consider the problem of personalized federated learning when there are known cluster structures within users. An intuitive approach would be to regularize the parameters so that…
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
Dake Zhang, Boxiang Lyu, Shuang Qiu +2
We study risk-sensitive reinforcement learning (RL), a crucial field due to its ability to enhance decision-making in scenarios where it is essential to manage uncertainty and mini…