Pretrained Encoders are All You Need
arXiv:2106.05139
Abstract
Data-efficiency and generalization are key challenges in deep learning and deep reinforcement learning as many models are trained on large-scale, domain-specific, and expensive-to-label datasets. Self-supervised models trained on large-scale uncurated datasets have shown successful transfer to diverse settings. We investigate using pretrained image representations and spatio-temporal attention for state representation learning in Atari. We also explore fine-tuning pretrained representations with self-supervised techniques, i.e., contrastive predictive coding, spatio-temporal contrastive learning, and augmentations. Our results show that pretrained representations are at par with state-of-the-art self-supervised methods trained on domain-specific data. Pretrained representations, thus, yield data and compute-efficient state representations. https://github.com/PAL-ML/PEARL_v1
References in corpus (9)
- Learning Transferable Visual Models From Natural Language Supervision
- Bootstrap your own latent: A new approach to self-supervised Learning
- Learning Representations by Maximizing Mutual Information Across Views
- DeepMind Control Suite
- MONet: Unsupervised Scene Decomposition and Representation
- Loss is its own Reward: Self-Supervision for Reinforcement Learning
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
- A Perspective on Objects and Systematic Generalization in Model-Based RL
- Personalizing Pre-trained Models