Publications (19)
Distributionally Robust Policy Learning under Concept Drifts
Jingyuan Wang, Zhimei Ren, Ruohan Zhan +1
Distributionally robust policy learning aims to find a policy that performs well under the worst-case distributional shift, and yet most existing methods for robust policy learning…
Fragility-aware Classification for Understanding Risk and Improving Generalization
Chen Yang, Zheng Cui, Daniel Zhuoyu Long +2
Classification models play a central role in data-driven decision-making applications such as medical diagnosis, recommendation systems, and risk assessment. Traditional performanc…
Confidence Intervals for Policy Evaluation in Adaptive Experiments
Vitor Hadad, David A. Hirshberg, Ruohan Zhan +2
Adaptive experiment designs can dramatically improve statistical efficiency in randomized trials, but they also complicate statistical inference. For example, it is now well known…
CT Image Reconstruction by Spatial-Radon Domain Data-Driven Tight Frame Regularization
Ruohan Zhan, Bin Dong
This paper proposes a spatial-Radon domain CT image reconstruction model based on data-driven tight frames (SRD-DDTF). The proposed SRD-DDTF model combines the idea of joint image…
Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits
Ruohan Zhan, Vitor Hadad, David A. Hirshberg +1
It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment as…
Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence
Zhiqi Zhang, Zhiyu Zeng, Ruohan Zhan +1
Randomized Controlled Trials (RCTs), or A/B testing, have become the gold standard for optimizing various operational policies on online platforms. However, RCTs on these platforms…
Policy Learning with Adaptively Collected Data
Ruohan Zhan, Zhimei Ren, Susan Athey +1
Learning optimal policies from historical data enables personalization in a wide variety of applications including healthcare, digital recommendations, and online education. The gr…
ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor
Wanqi Xue, Qingpeng Cai, Ruohan Zhan +4
Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dw…
Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities
Ruohan Zhan, Konstantina Christakopoulou, Ya Le +6
Most existing recommender systems focus primarily on matching users to content which maximizes user satisfaction on the platform. It is increasingly obvious, however, that content…
Adaptively Learning to Select-Rank in Online Platforms
Jingyuan Wang, Perry Dong, Ying Jin +2
Ranking algorithms are fundamental to various online platforms across e-commerce sites to content streaming services. Our research addresses the challenge of adaptively ranking ite…
Two-Stage Constrained Actor-Critic for Short Video Recommendation
Qingpeng Cai, Zhenghai Xue, Chi Zhang +9
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially intera…
Post Reinforcement Learning Inference
Vasilis Syrgkanis, Ruohan Zhan
We study estimation and inference using data collected by reinforcement learning (RL) algorithms. These algorithms adaptively experiment by interacting with individual units over m…
Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization
Sanath Kumar Krishnamurthy, Ruohan Zhan, Susan Athey +1
In many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That i…
Distortion Agnostic Deep Watermarking
Xiyang Luo, Ruohan Zhan, Huiwen Chang +2
Watermarking is the process of embedding information into an image that can survive under distortions, while requiring the encoded image to have little or no perceptual difference…
Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds
Yunbei Xu, Yuzhe Yuan, Ruohan Zhan
We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kull…
Deconfounding Duration Bias in Watch-time Prediction for Video Recommendation
Ruohan Zhan, Changhua Pei, Qiang Su +5
Watch-time prediction remains to be a key factor in reinforcing user engagement via video recommendations. It has become increasingly important given the ever-growing popularity of…
Constrained Reinforcement Learning for Short Video Recommendation
Qingpeng Cai, Ruohan Zhan, Chi Zhang +5
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users provide complex and…
Estimating Treatment Effects under Algorithmic Interference: A Structured Neural Networks Approach
Ruohan Zhan, Shichao Han, Yuchen Hu +1
Online user-generated content platforms allocate billions of dollars of promotional traffic through algorithms in two-sided marketplaces. To evaluate updates to these algorithms, p…
Statistical Properties of Robust Satisficing
Zhiyi Li, Yunbei Xu, Ruohan Zhan
The Robust Satisficing (RS) model is an emerging approach to robust optimization, offering streamlined procedures and robust generalization across various applications. However, th…