Masked Autoencoders As Spatiotemporal Learners
arXiv:2205.09113
Abstract
This paper studies a conceptually simple extension of Masked Autoencoders (MAE) to spatiotemporal representation learning from videos. We randomly mask out spacetime patches in videos and learn an autoencoder to reconstruct them in pixels. Interestingly, we show that our MAE method can learn strong representations with almost no inductive bias on spacetime (only except for patch and positional embeddings), and spacetime-agnostic random masking performs the best. We observe that the optimal masking ratio is as high as 90% (vs. 75% on images), supporting the hypothesis that this ratio is related to information redundancy of the data. A high masking ratio leads to a large speedup, e.g., > 4x in wall-clock time or even more. We report competitive results on several challenging video datasets using vanilla Vision Transformers. We observe that MAE can outperform supervised pre-training by large margins. We further report encouraging results of training on real-world, uncurated Instagram data. Our study suggests that the general framework of masked autoencoding (BERT, MAE, etc.) can be a unified methodology for representation learning with minimal domain knowledge.
Cited by in corpus (12)
- Knowledge Graph Self-Supervised Rationalization for Recommendation
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
- HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
- Masked Autoencoder for Self-Supervised Pre-training on Lidar Point Clouds
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology
- SVFAP: Self-supervised Video Facial Affect Perceiver
- Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement Learning
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
- Calibrated Self-supervised Vision Transformers Improve Intracranial Arterial Calcification Segmentation from Clinical CT Head Scans
- A Novel Tracking Framework for Devices in X-ray Leveraging Supplementary Cue-Driven Self-Supervised Features