5 papers
PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
Aishwarya Mandyam, Jason Meng, Ge Gao +4
Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
Aishwarya Mandyam, Shengpu Tang, Jiayu Yao +2
Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be e…
APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
Aishwarya Mandyam, Kalyani Limaye, Barbara E. Engelhardt +1
Off-policy evaluation (OPE) estimates the value of a contextual bandit policy prior to deployment. As such, OPE plays a critical role in ensuring safety in high-stakes domains such…
Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90
The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…
Kernel Density Bayesian Inverse Reinforcement Learning
Aishwarya Mandyam, Didong Li, Jiayu Yao +3
Inverse reinforcement learning (IRL) methods infer an agent's reward function using demonstrations of expert behavior. A Bayesian IRL approach models a distribution over candidate…