13 citations · 23 across the 7 of their papers we have counts for
12 papers
Evaluating for the long term: Learnings from industry
Leif Sigerson, Tom Cunningham, Winston Chou +22
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and shar…
Calibrated Recommendations with Contextual Bandits
Diego Feijer, Himan Abdollahpouri, Sanket Gupta +7
Spotify's Home page features a variety of content types, including music, podcasts, and audiobooks. However, historical data is heavily skewed toward music, making it challenging t…
Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration
Yan Chen, Qinxun Bai, Yiteng Zhang +4
Designing learning agents that explore efficiently in a complex environment has been widely recognized as a fundamental challenge in reinforcement learning. While a number of works…
Selectively Contextual Bandits
Claudia Roberts, Maria Dimakopoulou, Qifeng Qiao +2
Contextual bandits are widely used in industrial personalization systems. These online learning frameworks learn a treatment assignment policy in the presence of treatment effects…
Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy Learning
Aurélien Bibaut, Antoine Chambaz, Maria Dimakopoulou +2
Empirical risk minimization (ERM) is the workhorse of machine learning, whether for classification and regression or for off-policy policy learning, but its model-agnostic guarante…
Post-Contextual-Bandit Inference
Aurélien Bibaut, Antoine Chambaz, Maria Dimakopoulou +2
Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study particip…