collaborators

5 papers

cs.LG2026

PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data

Aishwarya Mandyam, Jason Meng, Ge Gao +4

Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…

cs.LG2026

CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

Aishwarya Mandyam, Shengpu Tang, Jiayu Yao +2

Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be e…

cs.LG2025

APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs

Aishwarya Mandyam, Kalyani Limaye, Barbara E. Engelhardt +1

Off-policy evaluation (OPE) estimates the value of a contextual bandit policy prior to deployment. As such, OPE plays a critical role in ensuring safety in high-stakes domains such…

cs.LG2025

Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025

Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90

The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…

cs.LG2025

Kernel Density Bayesian Inverse Reinforcement Learning

Aishwarya Mandyam, Didong Li, Jiayu Yao +3

Inverse reinforcement learning (IRL) methods infer an agent's reward function using demonstrations of expert behavior. A Bayesian IRL approach models a distribution over candidate…