activity
20152026
most citedPredictive Entropy Search for Bayesian Optimization with Unknown Constraints

107 citations · 130 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Latent-Constrained Conditional VAEs for Augmenting Large-Scale Climate Ensembles

Jacquelyn Shelton, Przemyslaw Polewski, Alexander Robel +2

Large climate-model ensembles are computationally expensive; yet many downstream analyses would benefit from additional, statistically consistent realizations of spatiotemporal cli…

cs.LG2022

Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach

Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6

Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…

cs.LG20214 cited

Regularized Behavior Value Estimation

Caglar Gulcehre, Sergio Gómez Colmenarejo, Ziyu Wang +7

Offline reinforcement learning restricts the learning process to rely only on logged-data without access to an environment. While this enables real-world applications, it also pose…

cs.LG2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

Caglar Gulcehre, Ziyu Wang, Alexander Novikov +15

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to lea…

cs.LG2019

Modular Meta-Learning with Shrinkage

Yutian Chen, Abram L. Friesen, Feryal Behbahani +4

Many real-world problems, including multi-speaker text-to-speech synthesis, can greatly benefit from the ability to meta-learn large models with only a few task-specific components…

cs.LG2018

One-Shot High-Fidelity Imitation: Training Large-Scale Deep Nets with RL

Tom Le Paine, Sergio Gómez Colmenarejo, Ziyu Wang +8

Humans are experts at high-fidelity imitation -- closely mimicking a demonstration, often in one attempt. Humans use this ability to quickly solve a task instance, and to bootstrap…