23 citations · 38 across the 11 of their papers we have counts for
12 papers · 1 filter
Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment
Yingchuan Sun, Shengpu Tang
Existing studies on reinforcement learning (RL) for sepsis management have mostly followed an established problem setup, in which patient data are aggregated into 4-hour time steps…
Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90
The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
Ziyan Wang, Zheng Wang, Xingwei Qu +5
Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs). Yet on-policy algorithms such as Group Relative Policy Optimization (GRPO) of…
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
Aishwarya Mandyam, Shengpu Tang, Jiayu Yao +2
Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be e…
Counterfactual-Augmented Importance Sampling for Semi-Offline Policy Evaluation
Shengpu Tang, Jenna Wiens
In applying reinforcement learning (RL) to high-stakes domains, quantitative and qualitative evaluation using observational data can help practitioners understand the generalizatio…
Leveraging Factored Action Spaces for Off-Policy Evaluation
Aaman Rebello, Shengpu Tang, Jenna Wiens +1
Off-policy evaluation (OPE) aims to estimate the benefit of following a counterfactual sequence of actions, given data collected from executed sequences. However, existing OPE esti…