activity
20112022
most citedOptimizing Dialogue Management with Reinforcement Learning: Experiments with the NJFun System

354 citations · 740 across the 11 of their papers we have counts for

collaborators
Showing cs.AIShow all

10 papers · 1 filter

cs.AI2017113 cited

Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning

Junhyuk Oh, Satinder Singh, Honglak Lee +1

As a step towards developing zero-shot task generalization capabilities in reinforcement learning (RL), we introduce a new RL problem where the agent should learn to execute sequen…

cs.AI2017

Repeated Inverse Reinforcement Learning

Kareem Amin, Nan Jiang, Satinder Singh

We introduce a novel repeated Inverse Reinforcement Learning problem: the agent has to act on behalf of a human in a sequence of tasks and wishes to minimize the number of tasks th…

cs.AI2017

Minimizing Maximum Regret in Commitment Constrained Sequential Decision Making

Qi Zhang, Satinder Singh, Edmund Durfee

In cooperative multiagent planning, it can often be beneficial for an agent to make commitments about aspects of its behavior to others, allowing them in turn to plan their own beh…

cs.AI2016

Control of Memory, Active Perception, and Action in Minecraft

Junhyuk Oh, Valliappa Chockalingam, Satinder Singh +1

In this paper, we introduce a new set of reinforcement learning (RL) tasks in Minecraft (a flexible 3D world). We then use these tasks to systematically compare and contrast existi…

cs.AI2016

Deep Learning for Reward Design to Improve Monte Carlo Tree Search in ATARI Games

Xiaoxiao Guo, Satinder Singh, Richard Lewis +1

Monte Carlo Tree Search (MCTS) methods have proven powerful in planning for sequential decision-making problems such as Go and video games, but their performance can be poor when t…

cs.AI201366 cited

Approximate Planning for Factored POMDPs using Belief State Simplification

David A. McAllester, Satinder Singh

We are interested in the problem of planning for factored POMDPs. Building on the recent results of Kearns, Mansour and Ng, we provide a planning algorithm for factored POMDPs that…