Showing cs.LGShow all
3 papers · 1 filter
cs.LG2023
Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits
Lequn Wang, Akshay Krishnamurthy, Aleksandrs Slivkins
We consider offline policy optimization (OPO) in contextual bandits, where one is given a fixed dataset of logged interactions. While pessimistic regularizers are typically used to…
cs.LG2021
Fairness of Exposure in Stochastic Bandits
Lequn Wang, Yiwei Bai, Wen Sun +1
Contextual bandit algorithms have become widely used for recommendation in online systems (e.g. marketplaces, music streaming, news), where they now wield substantial influence on…
cs.LG2018
CAB: Continuous Adaptive Blending Estimator for Policy Evaluation and Learning
Yi Su, Lequn Wang, Michele Santacatterina +1
The ability to perform offline A/B-testing and off-policy learning using logged contextual bandit feedback is highly desirable in a broad range of applications, including recommend…