activity
20172023
most citedScalable trust-region method for deep reinforcement learning using Kronecker-factored approximation

471 citations · 1.3k across the 24 of their papers we have counts for

collaborators
Showing cs.LGShow all

24 papers · 1 filter

cs.LG2023★ 55 cited

AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Yann Dubois, Xuechen Li, Rohan Taori +6

Large language models (LLMs) such as ChatGPT have seen widespread adoption due to their strong instruction-following abilities. Developing these LLMs involves a complex yet poorly…

cs.LG2023

TR0N: Translator Networks for 0-Shot Plug-and-Play Conditional Generation

Zhaoyan Liu, Noel Vouitsis, Satya Krishna Gorti +2

We propose TR0N, a highly general framework to turn pre-trained unconditional generative models, such as GANs and VAEs, into conditional models. The conditioning can be highly arbi…

cs.LG2022★ 3 cited

Exploring Low Rank Training of Deep Neural Networks

Siddhartha Rao Kamalakara, Acyr Locatelli, Bharat Venkitesh +3

Training deep neural networks in low rank, i.e. with factorised layers, is of particular interest to the community: it offers efficiency over unfactorised training in terms of both…

cs.LG2021

Learning Domain Invariant Representations in Goal-conditioned Block MDPs

Beining Han, Chongyi Zheng, Harris Chan +3

Deep Reinforcement Learning (RL) is successful in solving many complex Markov Decision Processes (MDPs) problems. However, agents often face unanticipated environmental changes aft…

cs.LG2020★ 3 cited

Planning from Pixels using Inverse Dynamics Models

Keiran Paster, Sheila A. McIlraith, Jimmy Ba

Learning task-agnostic dynamics models in high-dimensional observation spaces can be challenging for model-based RL agents. We propose a novel way to learn latent world models by l…

cs.LG2020★ 13 cited

A Study of Gradient Variance in Deep Learning

Fartash Faghri, David Duvenaud, David J. Fleet +1

The impact of gradient noise on training deep models is widely acknowledged but not well understood. In this context, we study the distribution of gradients during training. We int…