activity
20182022
most citedOff-policy Bandits with Deficient Support

23 citations · 26 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2022

Data-Driven Offline Decision-Making via Invariant Representation Learning

Han Qi, Yi Su, Aviral Kumar +1

The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active inte…

cs.IR20213 cited

Optimizing Rankings for Recommendation in Matching Markets

Yi Su, Magd Bayoumi, Thorsten Joachims

Based on the success of recommender systems in e-commerce, there is growing interest in their use in matching markets (e.g., labor). While this holds potential for improving market…

cs.LG202023 cited

Off-policy Bandits with Deficient Support

Noveen Sachdeva, Yi Su, Thorsten Joachims

Learning effective contextual-bandit policies from past actions of a deployed system is highly desirable in many settings (e.g. voice assistants, recommendation, search), since it…

cs.LG2020

Adaptive Estimator Selection for Off-Policy Evaluation

Yi Su, Pavithra Srinath, Akshay Krishnamurthy

We develop a generic data-driven method for estimator selection in off-policy policy evaluation settings. We establish a strong performance guarantee for the method, showing that i…

cs.LG2019

Doubly robust off-policy evaluation with shrinkage

Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy +1

We propose a new framework for designing estimators for off-policy evaluation in contextual bandits. Our approach is based on the asymptotically optimal doubly robust estimator, bu…

cs.LG2018

CAB: Continuous Adaptive Blending Estimator for Policy Evaluation and Learning

Yi Su, Lequn Wang, Michele Santacatterina +1

The ability to perform offline A/B-testing and off-policy learning using logged contextual bandit feedback is highly desirable in a broad range of applications, including recommend…