activity
20112022
most citedArtificial Intelligence and Life in 2030: The One Hundred Year Study on Artificial Intelligence

153 citations · 385 across the 25 of their papers we have counts for

collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2021

Adversarial Intrinsic Motivation for Reinforcement Learning

Ishan Durugkar, Mauricio Tec, Scott Niekum +1

Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we inve…

cs.LG2021

DEALIO: Data-Efficient Adversarial Learning for Imitation from Observation

Faraz Torabi, Garrett Warnell, Peter Stone

In imitation learning from observation IfO, a learning agent seeks to imitate a demonstrating agent using only observations of the demonstrated behavior without access to the contr…

cs.LG2021

Firefly Neural Architecture Descent: a General Approach for Growing Neural Networks

Lemeng Wu, Bo Liu, Peter Stone +1

We propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and archi…

cs.LG2020

Reinforcement Learning for Optimization of COVID-19 Mitigation policies

Varun Kompella, Roberto Capobianco, Stacy Jong +5

The year 2020 has seen the COVID-19 virus lead to one of the worst global pandemics in history. As a result, governments around the world are faced with the challenge of protecting…

cs.LG2020

Machine versus Human Attention in Deep Reinforcement Learning Tasks

Sihang Guo, Ruohan Zhang, Bo Liu +4

Deep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are…

cs.LG2020

Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy

Yunshu Du, Garrett Warnell, Assefaw Gebremedhin +2

Experience replay (ER) improves the data efficiency of off-policy reinforcement learning (RL) algorithms by allowing an agent to store and reuse its past experiences in a replay bu…