The Differentiable Cross-Entropy Method
arXiv:1909.12830
Abstract
We study the cross-entropy method (CEM) for the non-convex optimization of a continuous and parameterized objective function and introduce a differentiable variant that enables us to differentiate the output of CEM with respect to the objective function's parameters. In the machine learning setting this brings CEM inside of the end-to-end learning pipeline where this has otherwise been impossible. We show applications in a synthetic energy-based structured prediction task and in non-convex continuous control. In the control setting we show how to embed optimal action sequences into a lower-dimensional space. DCEM enables us to fine-tune CEM-based controllers with policy optimization.
ICML 2020
References in corpus (22)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- The NumPy array: a structure for efficient numerical computation
- DeepMind Control Suite
- Differentiable MPC for End-to-end Planning and Control
- Fast Context Adaptation via Meta-Learning
- Differentiable Convex Optimization Layers
- On Differentiating Parameterized Argmin and Argmax Problems with Application to Bi-level Optimization
- Monte Carlo Gradient Estimation in Machine Learning
- Exploring Model-based Planning with Policy Networks
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Deep Set Prediction Networks
- Adaptive and Safe Bayesian Optimization in High Dimensions via One-Dimensional Subspaces
- Variational Inference MPC for Bayesian Model-based Reinforcement Learning
- Design by adaptive sampling
- Learning Latent Plans from Play
- Objective Mismatch in Model-based Reinforcement Learning
- Disentangled State Space Representations
- The Limited Multi-Label Projection Layer
- Bayesian Optimization in Variational Latent Spaces with Dynamic Compression
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian Optimization
- Learning Convex Optimization Control Policies
- Prediction, Consistency, Curvature: Representation Learning for Locally-Linear Control