activity
20102020
most citedFrom Reinforcement Learning to Optimal Control: A unified framework for sequential decisions

9 citations · 17 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG20202 cited

On State Variables, Bandit Problems and POMDPs

Warren B Powell

State variables are easily the most subtle dimension of sequential decision problems. This is especially true in the context of active learning problems (bandit problems") where de…

math.OC20201 cited

Risk Directed Importance Sampling in Stochastic Dual Dynamic Programming with Hidden Markov Models for Grid Level Energy Storage

Joseph L. Durante, Juliana Nascimento, Warren B. Powell

Power systems that need to integrate renewables at a large scale must account for the high levels of uncertainty introduced by these power sources. This can be accomplished with a…

math.OC20205 cited

Reinforcement Learning via Parametric Cost Function Approximation for Multistage Stochastic Programming

Saeed Ghadimi, Raymond T. Perkins, Warren B. Powell

The most common approaches for solving stochastic resource allocation problems in the research literature is to either use value functions ("dynamic programming") or scenario trees…

cs.AI20199 cited

From Reinforcement Learning to Optimal Control: A unified framework for sequential decisions

Warren B Powell

There are over 15 distinct communities that work in the general area of sequential decisions and information, often referred to as decisions under uncertainty or stochastic optimiz…

stat.ML2016

Optimal Learning for Stochastic Optimization with Nonlinear Parametric Belief Models

Xinyu He, Warren B. Powell

We consider the problem of estimating the expected value of information (the knowledge gradient) for Bayesian learning problems where the belief model is nonlinear in the parameter…

math.OC2010

Stochastic Search with an Observable State Variable

Lauren A. Hannah, Warren B. Powell, David M. Blei

In this paper we study convex stochastic search problems where a noisy objective function value is observed after a decision is made. There are many stochastic search problems whos…