most citedAccelerating Quadratic Optimization with Reinforcement Learning

20 citations · 23 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL20211 cited

Discovering Non-monotonic Autoregressive Orderings with Variational Inference

Xuanlin Li, Brandon Trabucco, Dong Huk Park +4

The predominant approach for language modeling is to process sequences from left to right, but this eliminates a source of information: the order by which the sequence was generate…

cs.LG202120 cited

Accelerating Quadratic Optimization with Reinforcement Learning

Jeffrey Ichnowski, Paras Jain, Bartolomeo Stellato +6

First-order methods for quadratic optimization such as OSQP are widely used for large-scale machine learning and embedded optimal control, where many related problems must be rapid…

cs.RO2021

LazyDAgger: Reducing Context Switching in Interactive Imitation Learning

Ryan Hoque, Ashwin Balakrishna, Carl Putterman +6

Corrective interventions while a robot is learning to automate a task provide an intuitive method for a human supervisor to assist the robot and convey information about desired be…

cs.AI20201 cited

Connecting Context-specific Adaptation in Humans to Meta-learning

Rachit Dubey, Erin Grant, Michael Luo +2

Cognitive control, the ability of a system to adapt to the demands of a task, is an integral part of cognition. A widely accepted fact about cognitive control is that it is context…

cs.LG2020

RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem

Eric Liang, Zhanghao Wu, Michael Luo +3

Researchers and practitioners in the field of reinforcement learning (RL) frequently leverage parallel computation, which has led to a plethora of new algorithms and systems in the…

cs.LG20201 cited

IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks

Michael Luo, Jiahao Yao, Richard Liaw +2

The practical usage of reinforcement learning agents is often bottlenecked by the duration of training time. To accelerate training, practitioners often turn to distributed reinfor…