4 papers
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
Aishwarya Mandyam, Shengpu Tang, Jiayu Yao +2
Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be e…
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
Ziyan Wang, Zheng Wang, Xingwei Qu +5
Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs). Yet on-policy algorithms such as Group Relative Policy Optimization (GRPO) of…
Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment
Yingchuan Sun, Shengpu Tang
Existing studies on reinforcement learning (RL) for sepsis management have mostly followed an established problem setup, in which patient data are aggregated into 4-hour time steps…
Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90
The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…