Disentangled (Un)Controllable Features
arXiv:2211.00086 · doi:10.1109/SSCI52147.2023.10372006
Abstract
In the context of MDPs with high-dimensional states, downstream tasks are predominantly applied on a compressed, low-dimensional representation of the original input space. A variety of learning objectives have therefore been used to attain useful representations. However, these representations usually lack interpretability of the different features. We present a novel approach that is able to disentangle latent features into a controllable and an uncontrollable partition. We illustrate that the resulting partitioned representations are easily interpretable on three types of environments and show that, in a distribution of procedurally generated maze environments, it is feasible to interpretably employ a planning algorithm in the isolated controllable latent partition.
14 pages (8 main paper pages), 15 figures
References in corpus (9)
- Generative Adversarial Networks
- DeepMind Control Suite
- Contrastive Learning of Structured World Models
- Independently Controllable Factors
- Predictive Information Accelerates Learning in RL
- A Survey on Interpretable Reinforcement Learning
- Weakly Supervised Representation Learning with Sparse Perturbations
- Denoised MDPs: Learning World Models Better Than the World Itself
- Feature-Based Interpretable Reinforcement Learning based on State-Transition Models