Publications (98)
Problem Dependent Reinforcement Learning Bounds Which Can Identify Bandit Structure in MDPs
Andrea Zanette, Emma Brunskill
In order to make good decision under uncertainty an agent must learn from observations. To do so, two of the most common frameworks are Contextual Bandits and Markov Decision Proce…
Automated Reminders Reduce Incarceration for Missed Court Dates: Evidence from a Text Message Experiment
Alex Chohlas-Wood, Madison Coots, Joe Nudell +4
Millions of Americans must attend mandatory court dates every year. To boost appearance rates, jurisdictions nationwide are increasingly turning to automated reminders, but previou…
Supervised Pretraining Can Learn In-Context Reinforcement Learning
Jonathan N. Lee, Annie Xie, Aldo Pacchiano +4
Large transformer models trained on diverse datasets have shown a remarkable ability to learn in-context, achieving high few-shot performance on tasks they were not explicitly trai…
Design of Experiments for Stochastic Contextual Linear Bandits
Andrea Zanette, Kefan Dong, Jonathan Lee +1
In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…
Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds
Andrea Zanette, Emma Brunskill
Strong worst-case performance bounds for episodic reinforcement learning exist but fortunately in practice RL algorithms perform much better than such bounds would predict. Algorit…
Being Optimistic to Be Conservative: Quickly Learning a CVaR Policy
Ramtin Keramati, Christoph Dann, Alex Tamkin +1
While maximizing expected return is the goal in most reinforcement learning approaches, risk-sensitive objectives such as conditional value at risk (CVaR) are more suitable for man…
Policy Certificates: Towards Accountable Reinforcement Learning
Christoph Dann, Lihong Li, Wei Wei +1
The performance of a reinforcement learning algorithm can vary drastically during learning because of exploration. Existing algorithms provide little information about the quality…
Regret Bounds for Reinforcement Learning with Policy Advice
Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill
In some reinforcement learning problems an agent may be provided with a set of input policies, perhaps learned from prior experience or provided by advisors. We present a reinforce…
CORL: A Continuous-state Offset-dynamics Reinforcement Learner
Emma Brunskill, Bethany Leffler, Lihong Li +2
Continuous state spaces and stochastic, switching dynamics characterize a number of rich, realworld domains, such as robot navigation across varying terrain. We describe a reinforc…
Learning Abstract Models for Strategic Exploration and Fast Reward Transfer
Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri +4
Model-based reinforcement learning (RL) is appealing because (i) it enables planning and thus more strategic exploration, and (ii) by decoupling dynamics from rewards, it enables f…
Power Constrained Bandits
Jiayu Yao, Emma Brunskill, Weiwei Pan +2
Contextual bandits often provide simple and effective personalization in decision making problems, making them popular tools to deliver personalized interventions in mobile health…
Predicting Long Term Sequential Policy Value Using Softer Surrogates
Hyunji Nam, Allen Nie, Ge Gao +2
Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases…
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
Joy He-Yueya, Wanjing Anya Ma, Kanishk Gandhi +3
Language models (LMs) are increasingly used to simulate human-like responses in scenarios where accurately mimicking a population's behavior can guide decision-making, such as in d…
Learning When-to-Treat Policies
Xinkun Nie, Emma Brunskill, Stefan Wager
Many applied decision-making problems have a dynamic component: The policymaker needs not only to choose whom to treat, but also when to start which treatment. For example, a medic…
On Ensuring that Intelligent Machines Are Well-Behaved
Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto +1
Machine learning algorithms are everywhere, ranging from simple data analysis and pattern recognition tools used across the sciences to complex systems that achieve super-human per…
Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study
Ananya Bhattacharjee, Michael Liut, Matthew Jörke +3
Digital mental health (DMH) tools have extensively explored personalization of interventions to users' needs and contexts. However, this personalization often targets what support…
Improving Student Learning with Hybrid Human-AI Tutoring: A Three-Study Quasi-Experimental Investigation
Danielle R. Thomas, Jionghao Lin, Erin Gatz +8
Artificial intelligence (AI) applications to support human tutoring have potential to significantly improve learning outcomes, but engagement issues persist, especially among stude…
Frequentist Regret Bounds for Randomized Least-Squares Value Iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill +2
We consider the exploration-exploitation dilemma in finite-horizon reinforcement learning (RL). When the state space is large or continuous, traditional tabular approaches are unfe…
Sample Efficient Feature Selection for Factored MDPs
Zhaohan Daniel Guo, Emma Brunskill
In reinforcement learning, the state of the real world is often represented by feature vectors. However, not all of the features may be pertinent for solving the current task. We p…
Learning to be Fair: A Consequentialist Approach to Equitable Decision-Making
Alex Chohlas-Wood, Madison Coots, Henry Zhu +2
In an attempt to make algorithms fair, the machine learning literature has largely focused on equalizing decisions, outcomes, or error rates across race or gender groups. To illust…
Sublinear Optimal Policy Value Estimation in Contextual Bandits
Weihao Kong, Gregory Valiant, Emma Brunskill
We study the problem of estimating the expected reward of the optimal policy in the stochastic disjoint linear bandit setting. We prove that for certain settings it is possible to…
Bloom: Designing for LLM-Augmented Behavior Change Interactions
Matthew Jörke, Defne Genç, Valentin Teutschbein +7
Large language models (LLMs) offer novel opportunities to support health behavior change, yet existing work has narrowly focused on text-only interactions. Building on decades of H…
Identification of Subgroups With Similar Benefits in Off-Policy Policy Evaluation
Ramtin Keramati, Omer Gottesman, Leo Anthony Celi +2
Off-policy policy evaluation methods for sequential decision making can be used to help identify if a proposed decision policy is better than a current baseline policy. However, a…
Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding
Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky +1
When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluat…
Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs
Ashish Gurung, Ge Gao, Jordan Gutterman +6
Hybrid human-AI tutoring, where technology and humans jointly facilitate student learning, can be more beneficial than AI-only tutoring. However, preliminary evidence suggests that…
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
Allen Nie, Yash Chandak, Miroslav Suzara +6
Large language models (LLMs) are quickly being adopted in a wide range of learning experiences, especially via ubiquitous and broadly accessible chat interfaces like ChatGPT and Co…
Value Driven Representation for Human-in-the-Loop Reinforcement Learning
Ramtin Keramati, Emma Brunskill
Interactive adaptive systems powered by Reinforcement Learning (RL) have many potential applications, such as intelligent tutoring systems. In such systems there is typically an ex…
Learning Near Optimal Policies with Low Inherent Bellman Error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer +1
We study the exploration problem with approximate linear action-value functions in episodic reinforcement learning under the notion of low inherent Bellman error, a condition norma…
Importance Sampling with Unequal Support
Philip S. Thomas, Emma Brunskill
Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance samplin…
Separating value functions across time-scales
Joshua Romoff, Peter Henderson, Ahmed Touati +3
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to…
Estimating Optimal Policy Value in General Linear Contextual Bandits
Jonathan N. Lee, Weihao Kong, Aldo Pacchiano +2
In many bandit problems, the maximal reward achievable by a policy is often unknown in advance. We consider the problem of estimating the optimal policy value in the sublinear data…
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Peter Henderson, Jieru Hu, Joshua Romoff +3
Accurate reporting of energy and carbon usage is essential for understanding the potential climate impacts of machine learning research. We introduce a framework that makes this ea…
Provably Good Batch Reinforcement Learning Without Great Exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal +1
Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…
Incentive Decision Processes
Sashank J. Reddi, Emma Brunskill
We consider Incentive Decision Processes, where a principal seeks to reduce its costs due to another agent's behavior, by offering incentives to the agent for alternate behavior. W…
Universal Off-Policy Evaluation
Yash Chandak, Scott Niekum, Bruno Castro da Silva +3
When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…
PLOTS: Procedure Learning from Observations using Subtask Structure
Tong Mu, Karan Goel, Emma Brunskill
In many cases an intelligent agent may want to learn how to mimic a single observed demonstrated trajectory. In this work we consider how to perform such procedural learning from o…
Fast Exploration with Simplified Models and Approximately Optimistic Planning in Model Based Reinforcement Learning
Ramtin Keramati, Jay Whang, Patrick Cho +1
Humans learn to play video games significantly faster than the state-of-the-art reinforcement learning (RL) algorithms. People seem to build simple models that are easy to learn to…
Evaluating Treatment Prioritization Rules via Rank-Weighted Average Treatment Effects
Steve Yadlowsky, Scott Fleming, Nigam Shah +2
There are a number of available methods for selecting whom to prioritize for treatment, including ones based on treatment effect estimation, risk scoring, and hand-crafted rules. W…
A Statistical Test for the Benefits of Personalizing Interventions
Zhaoqi Li, Emma Brunskill
From medicine to marketing to social sciences, the promise of tailoring interventions to individuals is undeniable. However, practical applications force weighing personalization's…
Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling
Yao Liu, Pierre-Luc Bacon, Emma Brunskill
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods t…
The Online Coupon-Collector Problem and Its Application to Lifelong Reinforcement Learning
Emma Brunskill, Lihong Li
Transferring knowledge across a sequence of related tasks is an important challenge in reinforcement learning (RL). Despite much encouraging empirical evidence, there has been litt…
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets
Anirudhan Badrinath, Yannis Flet-Berliac, Allen Nie +1
Despite the recent advancements in offline reinforcement learning via supervised learning (RvS) and the success of the decision transformer (DT) architecture in various domains, DT…
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli +111
AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks.…
Online Model Selection for Reinforcement Learning with Function Approximation
Jonathan N. Lee, Aldo Pacchiano, Vidya Muthukumar +2
Deep reinforcement learning has achieved impressive successes yet often requires a very large amount of interaction data. This result is perhaps unsurprising, as using complicated…
Directed Exploration for Reinforcement Learning
Zhaohan Daniel Guo, Emma Brunskill
Efficient exploration is necessary to achieve good sample efficiency for reinforcement learning in general. From small, tabular settings such as gridworlds to large, continuous and…
Adaptive Interventions with User-Defined Goals for Health Behavior Change
Aishwarya Mandyam, Matthew Jörke, William Denton +2
Promoting healthy lifestyle behaviors remains a major public health concern, particularly due to their crucial role in preventing chronic conditions such as cancer, heart disease,…
Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
Andrea Zanette, Martin J. Wainwright, Emma Brunskill
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that…
Off-Policy Policy Gradient with State Distribution Correction
Yao Liu, Adith Swaminathan, Alekh Agarwal +1
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approac…
A PAC RL Algorithm for Episodic POMDPs
Zhaohan Daniel Guo, Shayan Doroudi, Emma Brunskill
Many interesting real world domains involve reinforcement learning (RL) in partially observable environments. Efficient learning in such domains is important, but existing sample c…
Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Commands
Thomas Kollar, Stefanie Tellex, Matthew Walter +8
Many task domains require robots to interpret and act upon natural language commands which are given by people and which refer to the robot's physical surroundings. Such interpreta…
Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors
Ryan Louie, Raj Sanjay Shah, Ifdita Hasan Orney +3
The growing demand for accessible mental health support requires training more counselors, yet existing approaches remain resource-intensive and difficult to scale. LLMs can realis…
Decoupling Learning Rules from Representations
Philip S. Thomas, Christoph Dann, Emma Brunskill
In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression…
Adaptive Instrument Design for Indirect Experiments
Yash Chandak, Shiv Shankar, Vasilis Syrgkanis +1
Indirect experiments provide a valuable framework for estimating treatment effects in situations where conducting randomized control trials (RCTs) is impractical or unethical. Unli…
Latent Contextual Bandits and their Application to Personalized Recommendations for New Users
Li Zhou, Emma Brunskill
Personalized recommendations for new users, also known as the cold-start problem, can be formulated as a contextual bandit problem. Existing contextual bandit algorithms generally…
Experiment Planning with Function Approximation
Aldo Pacchiano, Jonathan N. Lee, Emma Brunskill
We study the problem of experiment planning with function approximation in contextual bandit problems. In settings where there is a significant overhead to deploying adaptive algor…
Combining Parametric and Nonparametric Models for Off-Policy Evaluation
Omer Gottesman, Yao Liu, Scott Sussex +2
We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-pa…
Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study
Calvin Isley, Joshua Gilbert, Evangelos Kassos +9
While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instru…
Minimax-Regret Sample Selection in Randomized Experiments
Yuchen Hu, Henry Zhu, Emma Brunskill +1
Randomized controlled trials are often run in settings with many subpopulations that may have differential benefits from the treatment being evaluated. We consider the problem of s…
PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
Aishwarya Mandyam, Jason Meng, Ge Gao +4
Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…
Data-driven Error Estimation: Excess Risk Bounds without Class Complexity as Input
Sanath Kumar Krishnamurthy, Anna Lyubarskaja, Emma Brunskill +1
Constructing confidence intervals that are simultaneously valid across a class of estimates is central to tasks such as multiple mean estimation, generalization guarantees, and ada…
Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
Ryan Louie, Ananjan Nandi, William Fang +3
Recent works leverage LLMs to roleplay realistic social scenarios, aiding novices in practicing their social skills. However, simulating sensitive interactions, such as in mental h…
Short-Long Policy Evaluation with Novel Actions
Hyunji Alex Nam, Yash Chandak, Emma Brunskill
From incorporating LLMs in education, to identifying new drugs and improving ways to charge batteries, innovators constantly try new strategies in search of better long-term outcom…
Reinforcement Learning Tutor Better Supported Lower Performers in a Math Task
Sherry Ruan, Allen Nie, William Steenbergen +9
Resource limitations make it hard to provide all students with one of the most effective educational interventions: personalized instruction. Reinforcement learning could be a key…
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill
Human-designed reward functions for reinforcement learning (RL) agents are frequently misaligned with the humans' true, unobservable objectives, and thus act only as proxies. Optim…
Sample Complexity of Multi-task Reinforcement Learning
Emma Brunskill, Lihong Li
Transferring knowledge across a sequence of reinforcement-learning tasks is challenging, and has a number of important applications. Though there is encouraging empirical evidence…
Distilling Information from a Flood: A Possibility for the Use of Meta-Analysis and Systematic Review in Machine Learning Research
Peter Henderson, Emma Brunskill
The current flood of information in all areas of machine learning research, from computer vision to reinforcement learning, has made it difficult to make aggregate scientific infer…
Efficient Planning under Uncertainty with Macro-actions
Ruijie He, Emma Brunskill, Nicholas Roy
Deciding how to act in partially observable environments remains an active area of research. Identifying good sequences of decisions is particularly challenging when good control p…
Model-based Offline Reinforcement Learning with Local Misspecification
Kefan Dong, Yannis Flet-Berliac, Allen Nie +1
We present a model-based offline reinforcement learning policy performance lower bound that explicitly captures dynamics model misspecification and distribution mismatch and we pro…
Trading off rewards and errors in multi-armed bandits
Akram Erraqabi, Alessandro Lazaric, Michal Valko +2
In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…
GPTCoach: Towards LLM-Based Physical Activity Coaching
Matthew Jörke, Shardul Sapkota, Lyndsea Warkenthien +4
Mobile health applications show promise for scalable physical activity promotion but are often insufficiently personalized. In contrast, health coaching offers highly personalized…
RAPID: A Reachable Anytime Planner for Imprecisely-sensed Domains
Emma Brunskill, Stuart Russell
Despite the intractability of generic optimal partially observable Markov decision process planning, there exist important problems that have highly structured models. Previous res…
Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning
Christoph Dann, Tor Lattimore, Emma Brunskill
Statistical performance bounds for reinforcement learning (RL) algorithms can be critical for high-stakes applications like healthcare. This paper introduces a new framework for th…
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
Hyunji Nam, Omer Gottesman, Amy Zhang +3
Large language models (LLMs) built on existing reinforcement learning with human feedback (RLHF) frameworks typically optimize responses based on immediate turn-level human prefere…
Sample Complexity of Episodic Fixed-Horizon Reinforcement Learning
Christoph Dann, Emma Brunskill
Recently, there has been significant progress in understanding reinforcement learning in discounted infinite-horizon Markov decision processes (MDPs) by deriving tight sample compl…
When Simple Exploration is Sample Efficient: Identifying Sufficient Conditions for Random Exploration to Yield PAC RL Algorithms
Yao Liu, Emma Brunskill
Efficient exploration is one of the key challenges for reinforcement learning (RL) algorithms. Most traditional sample efficiency bounds require strategic exploration. Recently man…
Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization
Sanath Kumar Krishnamurthy, Ruohan Zhan, Susan Athey +1
In many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That i…
Using Options and Covariance Testing for Long Horizon Off-Policy Policy Evaluation
Zhaohan Daniel Guo, Philip S. Thomas, Emma Brunskill
Evaluating a policy by deploying it in the real world can be risky and costly. Off-policy policy evaluation (OPE) algorithms use historical data collected from running a previous p…
Policy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines
Philip S. Thomas, Emma Brunskill
We show how an action-dependent baseline can be used by the policy gradient theorem using function approximation, originally presented with action-independent baselines by (Sutton…
Sample Efficient Policy Search for Optimal Stopping Domains
Karan Goel, Christoph Dann, Emma Brunskill
Optimal stopping problems consider the question of deciding when to stop an observation-generating process in order to maximize a return. We examine the problem of simultaneously l…
Provably Efficient Reward-Agnostic Navigation with Linear Value Iteration
Andrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer +1
There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong as…
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
Stephane Hatgis-Kessell, Emma Brunskill
We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we replace classical RL algorith…
Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
Philip S. Thomas, Emma Brunskill
In this paper we present a new way of predicting the performance of a reinforcement learning policy given historical data that may have been generated by a different policy. The ab…
Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data
Allen Nie, Yannis Flet-Berliac, Deon R. Jordan +2
Offline reinforcement learning (RL) can be used to improve future performance by leveraging historical data. There exist many different algorithms for offline RL, and it is well re…
Giving Feedback on Interactive Student Programs with Meta-Exploration
Evan Zheran Liu, Moritz Stephan, Allen Nie +3
Developing interactive software, such as websites or games, is a particularly engaging way to learn computer science. However, teaching and giving feedback on such software is time…
Evaluating and Optimizing Educational Content with Large Language Model Judgments
Joy He-Yueya, Noah D. Goodman, Emma Brunskill
Creating effective educational materials generally requires expensive and time-consuming studies of student learning outcomes. To overcome this barrier, one idea is to build comput…
Oracle Inequalities for Model Selection in Offline Reinforcement Learning
Jonathan N. Lee, George Tucker, Ofir Nachum +2
In offline reinforcement learning (RL), a learner leverages prior logged data to learn a good policy without interacting with the environment. A major challenge in applying such me…
Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions
Omer Gottesman, Joseph Futoma, Yao Liu +4
Off-policy evaluation in reinforcement learning offers the chance of using observational data to improve future outcomes in domains such as healthcare and education, but safe deplo…
Offline Policy Optimization with Eligible Actions
Yao Liu, Yannis Flet-Berliac, Emma Brunskill
Offline policy optimization could have a large impact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling an…
Constraint Sampling Reinforcement Learning: Incorporating Expertise For Faster Learning
Tong Mu, Georgios Theocharous, David Arbour +1
Online reinforcement learning (RL) algorithms are often difficult to deploy in complex human-facing applications as they may learn slowly and have poor early performance. To addres…
Behaviour Policy Estimation in Off-Policy Policy Evaluation: Calibration Matters
Aniruddh Raghu, Omer Gottesman, Yao Liu +4
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empi…
GIANTS: Generative Insight Anticipation from Scientific Literature
Joy He-Yueya, Anikait Singh, Ge Gao +5
Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific discovery, their ability to per…
Play to Grade: Testing Coding Games as Classifying Markov Decision Process
Allen Nie, Emma Brunskill, Chris Piech
Contemporary coding education often presents students with the task of developing programs that have user interaction and complex dynamic systems, such as mouse based games. While…
Missingness as Stability: Understanding the Structure of Missingness in Longitudinal EHR data and its Impact on Reinforcement Learning in Healthcare
Scott L. Fleming, Kuhan Jeyapragasan, Tony Duan +4
There is an emerging trend in the reinforcement learning for healthcare literature. In order to prepare longitudinal, irregularly sampled, clinical datasets for reinforcement learn…
Sequential Transfer in Multi-armed Bandit with Finite Set of Models
Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill
Learning from prior tasks and transferring that experience to improve future performance is critical for building lifelong learning agents. Although results in supervised and reinf…
Active Learning for Stochastic Contextual Linear Bandits
Emma Brunskill, Ishani Karmarkar, Zhaoqi Li
A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions…
Predicting Long-Term Student Outcomes from Short-Term EdTech Log Data
Ge Gao, Amelia Leon, Andrea Jetten +4
Educational stakeholders are often particularly interested in sparse, delayed student outcomes, like end-of-year statewide exams. The rare occurrence of such assessments makes it h…
Representation Balancing MDPs for Off-Policy Policy Evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu +4
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value…
Online Stochastic Optimization under Correlated Bandit Feedback
Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill
In this paper we consider the problem of online stochastic optimization of a locally smooth function under bandit feedback. We introduce the high-confidence tree (HCT) algorithm, a…