activity
20152021
most citedBalancing Constraints and Rewards with Meta-Gradient D4PG

6 citations · 11 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20213 cited

Task-agnostic Continual Learning with Hybrid Probabilistic Models

Polina Kirichenko, Mehrdad Farajtabar, Dushyant Rao +6

Learning new tasks continuously without forgetting on a constantly changing data distribution is essential for real-world problems but extremely challenging for modern deep learnin…

cs.LG2020

Neural Rate Control for Video Encoding using Imitation Learning

Hongzi Mao, Chenjie Gu, Miaosen Wang +9

In modern video encoders, rate control is a critical component and has been heavily engineered. It decides how many bits to spend to encode each frame, in order to optimize the rat…

cs.LG20206 cited

Balancing Constraints and Rewards with Meta-Gradient D4PG

Dan A. Calian, Daniel J. Mankowitz, Tom Zahavy +4

Deploying Reinforcement Learning (RL) agents to solve real-world applications often requires satisfying complex system constraints. Often the constraint thresholds are incorrectly…

cs.LG2020

A maximum-entropy approach to off-policy evaluation in average-reward MDPs

Nevena Lazic, Dong Yin, Mehrdad Farajtabar +4

This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear…

cs.LG2020

An empirical investigation of the challenges of real-world reinforcement learning

Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz +4

Reinforcement learning (RL) has proven its worth in a series of artificial domains, and is beginning to show some successes in real-world scenarios. However, much of the research a…

cs.LG2019

Prediction, Consistency, Curvature: Representation Learning for Locally-Linear Control

Nir Levine, Yinlam Chow, Rui Shu +3

Many real-world sequential decision-making problems can be formulated as optimal control with high-dimensional observations and unknown dynamics. A promising approach is to embed t…