data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
arXiv:2202.03555
Abstract
While the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to general self-supervised learning, we present data2vec, a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture. Instead of predicting modality-specific targets such as words, visual tokens or units of human speech which are local in nature, data2vec predicts contextualized latent representations that contain information from the entire input. Experiments on the major benchmarks of speech recognition, image classification, and natural language understanding demonstrate a new state of the art or competitive performance to predominant approaches.
Cited by in corpus (22)
- DINOv2: Learning Robust Visual Features without Supervision
- Self-Supervised Speech Representation Learning: A Review
- AI and 6G into the Metaverse: Fundamentals, Challenges and Future Research Trends
- BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- Multimodal Neural Databases
- Image-Based Vehicle Classification by Synergizing Features from Supervised and Self-Supervised Learning Paradigms
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
- A Backdoor Approach with Inverted Labels Using Dirty Label-Flipping Attacks
- Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey)
- CrisisViT: A Robust Vision Transformer for Crisis Image Classification
- SSIN: Self-Supervised Learning for Rainfall Spatial Interpolation
- Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation
- An empirical study on speech restoration guided by self supervised speech representation
- Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning
- On convex decision regions in deep network representations
- SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
- Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
- Phone and speaker spatial organization in self-supervised speech representations
- Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners