4 papers
Dynamic Regret for Non-Stationary Linear Bandits via Misspecification Reductions
Zihao Hu, Yuan Yao, Jiheng Zhang +1
Many online decision-making problems involve both round-specific feasible actions and drifting reward models: eligible ad impressions, feasible prices, and available treatments can…
Learning to Bid with Unknown Private Values in Budget-Constrained First-Price Auctions
Zihao Hu, Yuxiao Wen, Yuan Yao +2
We study the operational problem of automated bidding in repeated first-price auctions under budget and return-on-spend (RoS) constraints. In this setting, an auto-bidder must tran…
Learning to Bid in Non-Stationary Repeated First-Price Auctions
Zihao Hu, Xiaoyu Fan, Yuan Yao +2
First-price auctions have recently gained significant traction in digital advertising markets, exemplified by Google's transition from second-price to first-price auctions. Unlike…
RL in Markov Games with Independent Function Approximation: Improved Sample Complexity Bound under the Local Access Model
Junyi Fan, Yuxuan Han, Jialin Zeng +4
Efficiently learning equilibria with large state and action spaces in general-sum Markov games while overcoming the curse of multi-agency is a challenging problem. Recent works hav…