6 papers
A Distribution Mapping Approach to Counterfactually Fair Reinforcement Learning
Jianhan Zhang, Jitao Wang, John D. Piette +3
Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcar…
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
Shijin Gong, Erhan Xu, Kai Ye +3
Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff betw…
Global Average Treatment Effects for Individualized Randomization Experiments with Aggregate Data
Shuguang Yu, Ting Li, Yuchen Lu +5
Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatme…
PyCFRL: A Python library for counterfactually fair offline reinforcement learning via sequential data preprocessing
Jianhan Zhang, Jitao Wang, Chengchun Shi +3
Reinforcement learning (RL) aims to learn and evaluate a sequential decision rule, often referred to as a "policy", that maximizes the population-level benefit in an environment ac…
Generalized Fitted Q-Iteration with Clustered Data
Liyuan Hu, Jitao Wang, Zhenke Wu +1
This paper focuses on reinforcement learning (RL) with clustered data, which is commonly encountered in healthcare applications. We propose a generalized fitted Q-iteration (FQI) a…
Counterfactually Fair Reinforcement Learning via Sequential Data Preprocessing
Jitao Wang, Chengchun Shi, John D. Piette +3
When applied in healthcare, reinforcement learning (RL) seeks to dynamically match the right interventions to subjects to maximize population benefit. However, the learned policy m…