activity
20192021
most citedKeeping Your Distance: Solving Sparse Reward Tasks Using Self-Balancing Shaped Rewards

21 citations · 45 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2021

Robustness Gym: Unifying the NLP Evaluation Landscape

Karan Goel, Nazneen Rajani, Jesse Vig +6

Despite impressive performance on standard benchmarks, deep neural networks are often brittle when deployed in real-world systems. Consequently, recent research has focused on test…

cs.CL2020

ESPRIT: Explaining Solutions to Physical Reasoning Tasks

Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan +7

Neural networks lack the ability to reason about qualitative physics and so cannot generalize to scenarios and tasks unseen during training. We propose ESPRIT, a framework for comm…

cs.AI201921 cited

Keeping Your Distance: Solving Sparse Reward Tasks Using Self-Balancing Shaped Rewards

Alexander Trott, Stephan Zheng, Caiming Xiong +1

While using shaped rewards can be beneficial when solving sparse reward tasks, their successful application often requires careful engineering and is problem specific. For instance…

cs.CL2019

Sketch-Fill-A-R: A Persona-Grounded Chit-Chat Generation Framework

Michael Shum, Stephan Zheng, Wojciech Kryściński +2

Human-like chit-chat conversation requires agents to generate responses that are fluent, engaging and consistent. We propose Sketch-Fill-A-R, a framework that uses a persona-memory…

cs.LG201910 cited

Learning World Graphs to Accelerate Hierarchical Reinforcement Learning

Wenling Shang, Alex Trott, Stephan Zheng +2

In many real-world scenarios, an autonomous agent often encounters various tasks within a single complex environment. We propose to build a graph abstraction over the environment s…

cs.LG201914 cited

On the Generalization Gap in Reparameterizable Reinforcement Learning

Huan Wang, Stephan Zheng, Caiming Xiong +1

Understanding generalization in reinforcement learning (RL) is a significant challenge, as many common assumptions of traditional supervised learning theory do not apply. We focus…