activity
20182022
most citedCombating the Compounding-Error Problem with a Multi-step Model

27 citations · 67 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2022

Towards Data-Driven Offline Simulations for Online Reinforcement Learning

Shengpu Tang, Felipe Vieira Frujeri, Dipendra Misra +4

Modern decision-making systems, from robots to web recommendation engines, are expected to adapt: to user preferences, changing circumstances or even new tasks. Yet, it is still un…

cs.LG20221 cited

Provable Safe Reinforcement Learning with Binary Feedback

Andrew Bennett, Dipendra Misra, Nathan Kallus

Safety is a crucial necessity in many applications of reinforcement learning (RL), whether robotic, automotive, or medical. Many existing approaches to safe RL rely on receiving nu…

cs.LG2022

Provably Sample-Efficient RL with Side Information about Latent Dynamics

Yao Liu, Dipendra Misra, Miro Dudík +1

We study reinforcement learning (RL) in settings where observations are high-dimensional, but where an RL agent has access to abstract knowledge about the structure of the state sp…

cs.LG202215 cited

Understanding Contrastive Learning Requires Incorporating Inductive Biases

Nikunj Saunshi, Jordan Ash, Surbhi Goel +5

Contrastive learning is a popular form of self-supervised learning that encourages augmentations (views) of the same input to have more similar representations compared to augmenta…

cs.LG202111 cited

Investigating the Role of Negatives in Contrastive Representation Learning

Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy +1

Noise contrastive learning is a popular technique for unsupervised representation learning. In this approach, a representation is obtained via reduction to supervised learning, whe…

cs.LG201913 cited

Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement Learning

Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy +1

We present an algorithm, HOMER, for exploration and reinforcement learning in rich observation environments that are summarizable by an unknown latent state space. The algorithm in…