papers

Publications (19)

cs.LG2025

Distributionally Robust Policy Learning under Concept Drifts

Jingyuan Wang, Zhimei Ren, Ruohan Zhan +1

Distributionally robust policy learning aims to find a policy that performs well under the worst-case distributional shift, and yet most existing methods for robust policy learning…

cs.LG2026

Fragility-aware Classification for Understanding Risk and Improving Generalization

Chen Yang, Zheng Cui, Daniel Zhuoyu Long +2

Classification models play a central role in data-driven decision-making applications such as medical diagnosis, recommendation systems, and risk assessment. Traditional performanc…

stat.ML2021

Confidence Intervals for Policy Evaluation in Adaptive Experiments

Vitor Hadad, David A. Hirshberg, Ruohan Zhan +2

Adaptive experiment designs can dramatically improve statistical efficiency in randomized trials, but they also complicate statistical inference. For example, it is now well known…

physics.med-ph2016

CT Image Reconstruction by Spatial-Radon Domain Data-Driven Tight Frame Regularization

Ruohan Zhan, Bin Dong

This paper proposes a spatial-Radon domain CT image reconstruction model based on data-driven tight frames (SRD-DDTF). The proposed SRD-DDTF model combines the idea of joint image…

stat.ML2021

Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits

Ruohan Zhan, Vitor Hadad, David A. Hirshberg +1

It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment as…

econ.EM2026

Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence

Zhiqi Zhang, Zhiyu Zeng, Ruohan Zhan +1

Randomized Controlled Trials (RCTs), or A/B testing, have become the gold standard for optimizing various operational policies on online platforms. However, RCTs on these platforms…

stat.ML2022

Policy Learning with Adaptively Collected Data

Ruohan Zhan, Zhimei Ren, Susan Athey +1

Learning optimal policies from historical data enables personalization in a wide variety of applications including healthcare, digital recommendations, and online education. The gr…

cs.IR2023

ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor

Wanqi Xue, Qingpeng Cai, Ruohan Zhan +4

Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dw…

cs.LG2021

Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities

Ruohan Zhan, Konstantina Christakopoulou, Ya Le +6

Most existing recommender systems focus primarily on matching users to content which maximizes user satisfaction on the platform. It is increasingly obvious, however, that content…

cs.LG2024

Adaptively Learning to Select-Rank in Online Platforms

Jingyuan Wang, Perry Dong, Ying Jin +2

Ranking algorithms are fundamental to various online platforms across e-commerce sites to content streaming services. Our research addresses the challenge of adaptively ranking ite…

cs.LG2024

Two-Stage Constrained Actor-Critic for Short Video Recommendation

Qingpeng Cai, Zhenghai Xue, Chi Zhang +9

The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially intera…

stat.ML2025

Post Reinforcement Learning Inference

Vasilis Syrgkanis, Ruohan Zhan

We study estimation and inference using data collected by reinforcement learning (RL) algorithms. These algorithms adaptively experiment by interacting with individual units over m…

cs.LG2023

Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization

Sanath Kumar Krishnamurthy, Ruohan Zhan, Susan Athey +1

In many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That i…

cs.MM2020

Distortion Agnostic Deep Watermarking

Xiyang Luo, Ruohan Zhan, Huiwen Chang +2

Watermarking is the process of embedding information into an image that can survive under distortions, while requiring the encoded image to have little or no perceptual difference…

cs.LG2026

Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds

Yunbei Xu, Yuzhe Yuan, Ruohan Zhan

We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kull…

cs.IR2022

Deconfounding Duration Bias in Watch-time Prediction for Video Recommendation

Ruohan Zhan, Changhua Pei, Qiang Su +5

Watch-time prediction remains to be a key factor in reinforcing user engagement via video recommendations. It has become increasingly important given the ever-growing popularity of…

cs.LG2022

Constrained Reinforcement Learning for Short Video Recommendation

Qingpeng Cai, Ruohan Zhan, Chi Zhang +5

The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users provide complex and…

econ.EM2026

Estimating Treatment Effects under Algorithmic Interference: A Structured Neural Networks Approach

Ruohan Zhan, Shichao Han, Yuchen Hu +1

Online user-generated content platforms allocate billions of dollars of promotional traffic through algorithms in two-sided marketplaces. To evaluate updates to these algorithms, p…

stat.ML2024

Statistical Properties of Robust Satisficing

Zhiyi Li, Yunbei Xu, Ruohan Zhan

The Robust Satisficing (RS) model is an emerging approach to robust optimization, offering streamlined procedures and robust generalization across various applications. However, th…