activity
20182026
most citedOff-policy Bandits with Deficient Support

23 citations · 28 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

LambdaRankIC: Directly Optimizing Rank IC for Financial Prediction

Yan Lin, Yihong Su, Yi Yang

In financial predictions, the performance of machine learning models is often assessed by Rank IC, which is the Spearman rank correlation between the model predictions and the real…

cs.LG2023★ 2 cited

Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective

Zeyu Zhang, Yi Su, Hui Yuan +5

Off-policy Learning to Rank (LTR) aims to optimize a ranker from data collected by a deployed logging policy. However, existing off-policy learning to rank methods often make stron…

cs.LG2022

Data-Driven Offline Decision-Making via Invariant Representation Learning

Han Qi, Yi Su, Aviral Kumar +1

The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active inte…

cs.LG2020★ 23 cited

Off-policy Bandits with Deficient Support

Noveen Sachdeva, Yi Su, Thorsten Joachims

Learning effective contextual-bandit policies from past actions of a deployed system is highly desirable in many settings (e.g. voice assistants, recommendation, search), since it…

cs.LG2020

Adaptive Estimator Selection for Off-Policy Evaluation

Yi Su, Pavithra Srinath, Akshay Krishnamurthy

We develop a generic data-driven method for estimator selection in off-policy policy evaluation settings. We establish a strong performance guarantee for the method, showing that i…

cs.LG2019

Doubly robust off-policy evaluation with shrinkage

Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy +1

We propose a new framework for designing estimators for off-policy evaluation in contextual bandits. Our approach is based on the asymptotically optimal doubly robust estimator, bu…