collaborators

6 papers

stat.ML2026

A Distribution Mapping Approach to Counterfactually Fair Reinforcement Learning

Jianhan Zhang, Jitao Wang, John D. Piette +3

Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcar…

cs.LG2026

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

Shijin Gong, Erhan Xu, Kai Ye +3

Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff betw…

stat.ME2026

Global Average Treatment Effects for Individualized Randomization Experiments with Aggregate Data

Shuguang Yu, Ting Li, Yuchen Lu +5

Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatme…

stat.ML2025

PyCFRL: A Python library for counterfactually fair offline reinforcement learning via sequential data preprocessing

Jianhan Zhang, Jitao Wang, Chengchun Shi +3

Reinforcement learning (RL) aims to learn and evaluate a sequential decision rule, often referred to as a "policy", that maximizes the population-level benefit in an environment ac…

cs.LG2025

Generalized Fitted Q-Iteration with Clustered Data

Liyuan Hu, Jitao Wang, Zhenke Wu +1

This paper focuses on reinforcement learning (RL) with clustered data, which is commonly encountered in healthcare applications. We propose a generalized fitted Q-iteration (FQI) a…

stat.ML2025

Counterfactually Fair Reinforcement Learning via Sequential Data Preprocessing

Jitao Wang, Chengchun Shi, John D. Piette +3

When applied in healthcare, reinforcement learning (RL) seeks to dynamically match the right interventions to subjects to maximize population benefit. However, the learned policy m…