papers

Publications (98)

cs.LG2019

Problem Dependent Reinforcement Learning Bounds Which Can Identify Bandit Structure in MDPs

Andrea Zanette, Emma Brunskill

In order to make good decision under uncertainty an agent must learn from observations. To do so, two of the most common frameworks are Contextual Bandits and Markov Decision Proce…

stat.AP2024

Automated Reminders Reduce Incarceration for Missed Court Dates: Evidence from a Text Message Experiment

Alex Chohlas-Wood, Madison Coots, Joe Nudell +4

Millions of Americans must attend mandatory court dates every year. To boost appearance rates, jurisdictions nationwide are increasingly turning to automated reminders, but previou…

cs.LG2023

Supervised Pretraining Can Learn In-Context Reinforcement Learning

Jonathan N. Lee, Annie Xie, Aldo Pacchiano +4

Large transformer models trained on diverse datasets have shown a remarkable ability to learn in-context, achieving high few-shot performance on tasks they were not explicitly trai…

cs.LG2021

Design of Experiments for Stochastic Contextual Linear Bandits

Andrea Zanette, Kefan Dong, Jonathan Lee +1

In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…

cs.LG2019

Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds

Andrea Zanette, Emma Brunskill

Strong worst-case performance bounds for episodic reinforcement learning exist but fortunately in practice RL algorithms perform much better than such bounds would predict. Algorit…

cs.LG2020

Being Optimistic to Be Conservative: Quickly Learning a CVaR Policy

Ramtin Keramati, Christoph Dann, Alex Tamkin +1

While maximizing expected return is the goal in most reinforcement learning approaches, risk-sensitive objectives such as conditional value at risk (CVaR) are more suitable for man…

cs.LG2019

Policy Certificates: Towards Accountable Reinforcement Learning

Christoph Dann, Lihong Li, Wei Wei +1

The performance of a reinforcement learning algorithm can vary drastically during learning because of exploration. Existing algorithms provide little information about the quality…

stat.ML2013

Regret Bounds for Reinforcement Learning with Policy Advice

Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill

In some reinforcement learning problems an agent may be provided with a set of input policies, perhaps learned from prior experience or provided by advisors. We present a reinforce…

cs.LG2012

CORL: A Continuous-state Offset-dynamics Reinforcement Learner

Emma Brunskill, Bethany Leffler, Lihong Li +2

Continuous state spaces and stochastic, switching dynamics characterize a number of rich, realworld domains, such as robot navigation across varying terrain. We describe a reinforc…

cs.LG2020

Learning Abstract Models for Strategic Exploration and Fast Reward Transfer

Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri +4

Model-based reinforcement learning (RL) is appealing because (i) it enables planning and thus more strategic exploration, and (ii) by decoupling dynamics from rewards, it enables f…

cs.LG2021

Power Constrained Bandits

Jiayu Yao, Emma Brunskill, Weiwei Pan +2

Contextual bandits often provide simple and effective personalization in decision making problems, making them popular tools to deliver personalized interventions in mobile health…

cs.AI2025

Predicting Long Term Sequential Policy Value Using Softer Surrogates

Hyunji Nam, Allen Nie, Ge Gao +2

Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases…

cs.CL2024

Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models

Joy He-Yueya, Wanjing Anya Ma, Kanishk Gandhi +3

Language models (LMs) are increasingly used to simulate human-like responses in scenarios where accurately mimicking a population's behavior can guide decision-making, such as in d…

stat.ME2020

Learning When-to-Treat Policies

Xinkun Nie, Emma Brunskill, Stefan Wager

Many applied decision-making problems have a dynamic component: The policymaker needs not only to choose whom to treat, but also when to start which treatment. For example, a medic…

cs.AI2017

On Ensuring that Intelligent Machines Are Well-Behaved

Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto +1

Machine learning algorithms are everywhere, ranging from simple data analysis and pattern recognition tools used across the sciences to complex systems that achieve super-human per…

cs.HC2026

Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study

Ananya Bhattacharjee, Michael Liut, Matthew Jörke +3

Digital mental health (DMH) tools have extensively explored personalization of interventions to users' needs and contexts. However, this personalization often targets what support…

cs.HC2023

Improving Student Learning with Hybrid Human-AI Tutoring: A Three-Study Quasi-Experimental Investigation

Danielle R. Thomas, Jionghao Lin, Erin Gatz +8

Artificial intelligence (AI) applications to support human tutoring have potential to significantly improve learning outcomes, but engagement issues persist, especially among stude…

cs.LG2023

Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

Andrea Zanette, David Brandfonbrener, Emma Brunskill +2

We consider the exploration-exploitation dilemma in finite-horizon reinforcement learning (RL). When the state space is large or continuous, traditional tabular approaches are unfe…

cs.LG2017

Sample Efficient Feature Selection for Factored MDPs

Zhaohan Daniel Guo, Emma Brunskill

In reinforcement learning, the state of the real world is often represented by feature vectors. However, not all of the features may be pertinent for solving the current task. We p…

cs.LG2024

Learning to be Fair: A Consequentialist Approach to Equitable Decision-Making

Alex Chohlas-Wood, Madison Coots, Henry Zhu +2

In an attempt to make algorithms fair, the machine learning literature has largely focused on equalizing decisions, outcomes, or error rates across race or gender groups. To illust…

cs.LG2019

Sublinear Optimal Policy Value Estimation in Contextual Bandits

Weihao Kong, Gregory Valiant, Emma Brunskill

We study the problem of estimating the expected reward of the optimal policy in the stochastic disjoint linear bandit setting. We prove that for certain settings it is possible to…

cs.HC2026

Bloom: Designing for LLM-Augmented Behavior Change Interactions

Matthew Jörke, Defne Genç, Valentin Teutschbein +7

Large language models (LLMs) offer novel opportunities to support health behavior change, yet existing work has narrowly focused on text-only interactions. Building on decades of H…

cs.LG2021

Identification of Subgroups With Similar Benefits in Off-Policy Policy Evaluation

Ramtin Keramati, Omer Gottesman, Leo Anthony Celi +2

Off-policy policy evaluation methods for sequential decision making can be used to help identify if a proposed decision policy is better than a current baseline policy. However, a…

stat.ML2020

Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding

Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky +1

When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluat…

cs.CY2026

Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs

Ashish Gurung, Ge Gao, Jordan Gutterman +6

Hybrid human-AI tutoring, where technology and humans jointly facilitate student learning, can be more beneficial than AI-only tutoring. However, preliminary evidence suggests that…

cs.CY2025

The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances

Allen Nie, Yash Chandak, Miroslav Suzara +6

Large language models (LLMs) are quickly being adopted in a wide range of learning experiences, especially via ubiquitous and broadly accessible chat interfaces like ChatGPT and Co…

cs.AI2020

Value Driven Representation for Human-in-the-Loop Reinforcement Learning

Ramtin Keramati, Emma Brunskill

Interactive adaptive systems powered by Reinforcement Learning (RL) have many potential applications, such as intelligent tutoring systems. In such systems there is typically an ex…

cs.LG2020

Learning Near Optimal Policies with Low Inherent Bellman Error

Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer +1

We study the exploration problem with approximate linear action-value functions in episodic reinforcement learning under the notion of low inherent Bellman error, a condition norma…

cs.LG2016

Importance Sampling with Unequal Support

Philip S. Thomas, Emma Brunskill

Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance samplin…

cs.LG2019

Separating value functions across time-scales

Joshua Romoff, Peter Henderson, Ahmed Touati +3

In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to…

cs.LG2023

Estimating Optimal Policy Value in General Linear Contextual Bandits

Jonathan N. Lee, Weihao Kong, Aldo Pacchiano +2

In many bandit problems, the maximal reward achievable by a policy is often unknown in advance. We consider the problem of estimating the optimal policy value in the sublinear data…

cs.CY2022

Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning

Peter Henderson, Jieru Hu, Joshua Romoff +3

Accurate reporting of energy and carbon usage is essential for understanding the potential climate impacts of machine learning research. We introduce a framework that makes this ea…

cs.LG2020

Provably Good Batch Reinforcement Learning Without Great Exploration

Yao Liu, Adith Swaminathan, Alekh Agarwal +1

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…

cs.GT2012

Incentive Decision Processes

Sashank J. Reddi, Emma Brunskill

We consider Incentive Decision Processes, where a principal seeks to reduce its costs due to another agent's behavior, by offering incentives to the agent for alternate behavior. W…

cs.LG2021

Universal Off-Policy Evaluation

Yash Chandak, Scott Niekum, Bruno Castro da Silva +3

When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…

cs.LG2019

PLOTS: Procedure Learning from Observations using Subtask Structure

Tong Mu, Karan Goel, Emma Brunskill

In many cases an intelligent agent may want to learn how to mimic a single observed demonstrated trajectory. In this work we consider how to perform such procedural learning from o…

cs.AI2018

Fast Exploration with Simplified Models and Approximately Optimistic Planning in Model Based Reinforcement Learning

Ramtin Keramati, Jay Whang, Patrick Cho +1

Humans learn to play video games significantly faster than the state-of-the-art reinforcement learning (RL) algorithms. People seem to build simple models that are easy to learn to…

stat.ME2023

Evaluating Treatment Prioritization Rules via Rank-Weighted Average Treatment Effects

Steve Yadlowsky, Scott Fleming, Nigam Shah +2

There are a number of available methods for selecting whom to prioritize for treatment, including ones based on treatment effect estimation, risk scoring, and hand-crafted rules. W…

stat.ME2026

A Statistical Test for the Benefits of Personalizing Interventions

Zhaoqi Li, Emma Brunskill

From medicine to marketing to social sciences, the promise of tailoring interventions to individuals is undeniable. However, practical applications force weighing personalization's…

cs.LG2020

Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling

Yao Liu, Pierre-Luc Bacon, Emma Brunskill

Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods t…

cs.LG2015

The Online Coupon-Collector Problem and Its Application to Lifelong Reinforcement Learning

Emma Brunskill, Lihong Li

Transferring knowledge across a sequence of related tasks is an important challenge in reinforcement learning (RL). Despite much encouraging empirical evidence, there has been litt…

cs.LG2023

Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets

Anirudhan Badrinath, Yannis Flet-Berliac, Allen Nie +1

Despite the recent advancements in offline reinforcement learning via supervised learning (RvS) and the success of the decision transformer (DT) architecture in various domains, DT…

cs.LG2022

On the Opportunities and Risks of Foundation Models

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli +111

AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks.…

cs.LG2020

Online Model Selection for Reinforcement Learning with Function Approximation

Jonathan N. Lee, Aldo Pacchiano, Vidya Muthukumar +2

Deep reinforcement learning has achieved impressive successes yet often requires a very large amount of interaction data. This result is perhaps unsurprising, as using complicated…

cs.LG2019

Directed Exploration for Reinforcement Learning

Zhaohan Daniel Guo, Emma Brunskill

Efficient exploration is necessary to achieve good sample efficiency for reinforcement learning in general. From small, tabular settings such as gridworlds to large, continuous and…

cs.LG2024

Adaptive Interventions with User-Defined Goals for Health Behavior Change

Aishwarya Mandyam, Matthew Jörke, William Denton +2

Promoting healthy lifestyle behaviors remains a major public health concern, particularly due to their crucial role in preventing chronic conditions such as cancer, heart disease,…

cs.LG2021

Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning

Andrea Zanette, Martin J. Wainwright, Emma Brunskill

Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that…

cs.LG2019

Off-Policy Policy Gradient with State Distribution Correction

Yao Liu, Adith Swaminathan, Alekh Agarwal +1

We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approac…

cs.LG2016

A PAC RL Algorithm for Episodic POMDPs

Zhaohan Daniel Guo, Shayan Doroudi, Emma Brunskill

Many interesting real world domains involve reinforcement learning (RL) in partially observable environments. Efficient learning in such domains is important, but existing sample c…

cs.CL2017

Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Commands

Thomas Kollar, Stefanie Tellex, Matthew Walter +8

Many task domains require robots to interpret and act upon natural language commands which are given by people and which refer to the robot's physical surroundings. Such interpreta…

cs.HC2026

Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors

Ryan Louie, Raj Sanjay Shah, Ifdita Hasan Orney +3

The growing demand for accessible mental health support requires training more counselors, yet existing approaches remain resource-intensive and difficult to scale. LLMs can realis…

cs.AI2017

Decoupling Learning Rules from Representations

Philip S. Thomas, Christoph Dann, Emma Brunskill

In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression…

cs.LG2023

Adaptive Instrument Design for Indirect Experiments

Yash Chandak, Shiv Shankar, Vasilis Syrgkanis +1

Indirect experiments provide a valuable framework for estimating treatment effects in situations where conducting randomized control trials (RCTs) is impractical or unethical. Unli…

cs.LG2016

Latent Contextual Bandits and their Application to Personalized Recommendations for New Users

Li Zhou, Emma Brunskill

Personalized recommendations for new users, also known as the cold-start problem, can be formulated as a contextual bandit problem. Existing contextual bandit algorithms generally…

cs.LG2024

Experiment Planning with Function Approximation

Aldo Pacchiano, Jonathan N. Lee, Emma Brunskill

We study the problem of experiment planning with function approximation in contextual bandit problems. In settings where there is a significant overhead to deploying adaptive algor…

cs.LG2019

Combining Parametric and Nonparametric Models for Off-Policy Evaluation

Omer Gottesman, Yao Liu, Scott Sussex +2

We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-pa…

cs.CY2025

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

Calvin Isley, Joshua Gilbert, Evangelos Kassos +9

While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instru…

stat.ME2024

Minimax-Regret Sample Selection in Randomized Experiments

Yuchen Hu, Henry Zhu, Emma Brunskill +1

Randomized controlled trials are often run in settings with many subpopulations that may have differential benefits from the treatment being evaluated. We consider the problem of s…

cs.LG2026

PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data

Aishwarya Mandyam, Jason Meng, Ge Gao +4

Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…

cs.LG2026

Data-driven Error Estimation: Excess Risk Bounds without Class Complexity as Input

Sanath Kumar Krishnamurthy, Anna Lyubarskaja, Emma Brunskill +1

Constructing confidence intervals that are simultaneously valid across a class of estimates is central to tasks such as multiple mean estimation, generalization guarantees, and ada…

cs.CL2024

Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles

Ryan Louie, Ananjan Nandi, William Fang +3

Recent works leverage LLMs to roleplay realistic social scenarios, aiding novices in practicing their social skills. However, simulating sensitive interactions, such as in mental h…

cs.LG2024

Short-Long Policy Evaluation with Novel Actions

Hyunji Alex Nam, Yash Chandak, Emma Brunskill

From incorporating LLMs in education, to identifying new drugs and improving ways to charge batteries, innovators constantly try new strategies in search of better long-term outcom…

cs.AI2023

Reinforcement Learning Tutor Better Supported Lower Performers in a Math Task

Sherry Ruan, Allen Nie, William Steenbergen +9

Resource limitations make it hard to provide all students with one of the most effective educational interventions: personalized instruction. Reinforcement learning could be a key…

cs.AI2026

Repairing Reward Functions with Feedback to Mitigate Reward Hacking

Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill

Human-designed reward functions for reinforcement learning (RL) agents are frequently misaligned with the humans' true, unobservable objectives, and thus act only as proxies. Optim…

cs.LG2013

Sample Complexity of Multi-task Reinforcement Learning

Emma Brunskill, Lihong Li

Transferring knowledge across a sequence of reinforcement-learning tasks is challenging, and has a number of important applications. Though there is encouraging empirical evidence…

cs.DL2018

Distilling Information from a Flood: A Possibility for the Use of Meta-Analysis and Systematic Review in Machine Learning Research

Peter Henderson, Emma Brunskill

The current flood of information in all areas of machine learning research, from computer vision to reinforcement learning, has made it difficult to make aggregate scientific infer…

cs.AI2014

Efficient Planning under Uncertainty with Macro-actions

Ruijie He, Emma Brunskill, Nicholas Roy

Deciding how to act in partially observable environments remains an active area of research. Identifying good sequences of decisions is particularly challenging when good control p…

cs.LG2023

Model-based Offline Reinforcement Learning with Local Misspecification

Kefan Dong, Yannis Flet-Berliac, Allen Nie +1

We present a model-based offline reinforcement learning policy performance lower bound that explicitly captures dynamics model misspecification and distribution mismatch and we pro…

cs.LG2026

Trading off rewards and errors in multi-armed bandits

Akram Erraqabi, Alessandro Lazaric, Michal Valko +2

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…

cs.HC2025

GPTCoach: Towards LLM-Based Physical Activity Coaching

Matthew Jörke, Shardul Sapkota, Lyndsea Warkenthien +4

Mobile health applications show promise for scalable physical activity promotion but are often insufficiently personalized. In contrast, health coaching offers highly personalized…

cs.AI2012

RAPID: A Reachable Anytime Planner for Imprecisely-sensed Domains

Emma Brunskill, Stuart Russell

Despite the intractability of generic optimal partially observable Markov decision process planning, there exist important problems that have highly structured models. Previous res…

cs.LG2018

Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning

Christoph Dann, Tor Lattimore, Emma Brunskill

Statistical performance bounds for reinforcement learning (RL) algorithms can be critical for high-stakes applications like healthcare. This paper introduces a new framework for th…

cs.CL2025

Efficient RL for optimizing conversation level outcomes with an LLM-based tutor

Hyunji Nam, Omer Gottesman, Amy Zhang +3

Large language models (LLMs) built on existing reinforcement learning with human feedback (RLHF) frameworks typically optimize responses based on immediate turn-level human prefere…

stat.ML2016

Sample Complexity of Episodic Fixed-Horizon Reinforcement Learning

Christoph Dann, Emma Brunskill

Recently, there has been significant progress in understanding reinforcement learning in discounted infinite-horizon Markov decision processes (MDPs) by deriving tight sample compl…

cs.LG2019

When Simple Exploration is Sample Efficient: Identifying Sufficient Conditions for Random Exploration to Yield PAC RL Algorithms

Yao Liu, Emma Brunskill

Efficient exploration is one of the key challenges for reinforcement learning (RL) algorithms. Most traditional sample efficiency bounds require strategic exploration. Recently man…

cs.LG2023

Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization

Sanath Kumar Krishnamurthy, Ruohan Zhan, Susan Athey +1

In many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That i…

cs.AI2017

Using Options and Covariance Testing for Long Horizon Off-Policy Policy Evaluation

Zhaohan Daniel Guo, Philip S. Thomas, Emma Brunskill

Evaluating a policy by deploying it in the real world can be risky and costly. Off-policy policy evaluation (OPE) algorithms use historical data collected from running a previous p…

cs.AI2017

Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

Philip S. Thomas, Emma Brunskill

We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton…

cs.AI2017

Sample Efficient Policy Search for Optimal Stopping Domains

Karan Goel, Christoph Dann, Emma Brunskill

Optimal stopping problems consider the question of deciding when to stop an observation-generating process in order to maximize a return. We examine the problem of simultaneously l…

cs.LG2020

Provably Efficient Reward-Agnostic Navigation with Linear Value Iteration

Andrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer +1

There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong as…

cs.LG2026

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

Stephane Hatgis-Kessell, Emma Brunskill

We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we replace classical RL algorith…

cs.LG2016

Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning

Philip S. Thomas, Emma Brunskill

In this paper we present a new way of predicting the performance of a reinforcement learning policy given historical data that may have been generated by a different policy. The ab…

cs.LG2023

Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data

Allen Nie, Yannis Flet-Berliac, Deon R. Jordan +2

Offline reinforcement learning (RL) can be used to improve future performance by leveraging historical data. There exist many different algorithms for offline RL, and it is well re…

cs.LG2022

Giving Feedback on Interactive Student Programs with Meta-Exploration

Evan Zheran Liu, Moritz Stephan, Allen Nie +3

Developing interactive software, such as websites or games, is a particularly engaging way to learn computer science. However, teaching and giving feedback on such software is time…

cs.AI2024

Evaluating and Optimizing Educational Content with Large Language Model Judgments

Joy He-Yueya, Noah D. Goodman, Emma Brunskill

Creating effective educational materials generally requires expensive and time-consuming studies of student learning outcomes. To overcome this barrier, one idea is to build comput…

cs.LG2022

Oracle Inequalities for Model Selection in Offline Reinforcement Learning

Jonathan N. Lee, George Tucker, Ofir Nachum +2

In offline reinforcement learning (RL), a learner leverages prior logged data to learn a good policy without interacting with the environment. A major challenge in applying such me…

cs.LG2020

Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions

Omer Gottesman, Joseph Futoma, Yao Liu +4

Off-policy evaluation in reinforcement learning offers the chance of using observational data to improve future outcomes in domains such as healthcare and education, but safe deplo…

cs.LG2022

Offline Policy Optimization with Eligible Actions

Yao Liu, Yannis Flet-Berliac, Emma Brunskill

Offline policy optimization could have a large impact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling an…

cs.AI2021

Constraint Sampling Reinforcement Learning: Incorporating Expertise For Faster Learning

Tong Mu, Georgios Theocharous, David Arbour +1

Online reinforcement learning (RL) algorithms are often difficult to deploy in complex human-facing applications as they may learn slowly and have poor early performance. To addres…

cs.LG2018

Behaviour Policy Estimation in Off-Policy Policy Evaluation: Calibration Matters

Aniruddh Raghu, Omer Gottesman, Yao Liu +4

In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empi…

cs.CL2026

GIANTS: Generative Insight Anticipation from Scientific Literature

Joy He-Yueya, Anikait Singh, Ge Gao +5

Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific discovery, their ability to per…

cs.AI2021

Play to Grade: Testing Coding Games as Classifying Markov Decision Process

Allen Nie, Emma Brunskill, Chris Piech

Contemporary coding education often presents students with the task of developing programs that have user interaction and complex dynamic systems, such as mouse based games. While…

cs.LG2019

Missingness as Stability: Understanding the Structure of Missingness in Longitudinal EHR data and its Impact on Reinforcement Learning in Healthcare

Scott L. Fleming, Kuhan Jeyapragasan, Tony Duan +4

There is an emerging trend in the reinforcement learning for healthcare literature. In order to prepare longitudinal, irregularly sampled, clinical datasets for reinforcement learn…

stat.ML2013

Sequential Transfer in Multi-armed Bandit with Finite Set of Models

Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill

Learning from prior tasks and transferring that experience to improve future performance is critical for building lifelong learning agents. Although results in supervised and reinf…

cs.LG2026

Active Learning for Stochastic Contextual Linear Bandits

Emma Brunskill, Ishani Karmarkar, Zhaoqi Li

A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions…

cs.CY2025

Predicting Long-Term Student Outcomes from Short-Term EdTech Log Data

Ge Gao, Amelia Leon, Andrea Jetten +4

Educational stakeholders are often particularly interested in sparse, delayed student outcomes, like end-of-year statewide exams. The rare occurrence of such assessments makes it h…

cs.LG2019

Representation Balancing MDPs for Off-Policy Policy Evaluation

Yao Liu, Omer Gottesman, Aniruddh Raghu +4

We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value…

stat.ML2014

Online Stochastic Optimization under Correlated Bandit Feedback

Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill

In this paper we consider the problem of online stochastic optimization of a locally smooth function under bandit feedback. We introduce the high-confidence tree (HCT) algorithm, a…