4 papers
Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning
Xiangkun Wu, Qianglin Wen, Yingying Zhang +3
A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequenti…
A Two-armed Bandit Framework for A/B Testing
Jinjuan Wang, Qianglin Wen, Yu Zhang +2
A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy again…
Deep Distributional Learning with Non-crossing Quantile Network
Guohao Shen, Runpeng Dai, Guojun Wu +3
In this paper, we introduce a non-crossing quantile (NQ) network for conditional distribution learning. By leveraging non-negative activation functions, the NQ network ensures that…
Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
Shuguang Yu, Shuxing Fang, Ruixin Peng +3
This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data liter…