OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
arXiv:2202.03052
Abstract
In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.
Accepted at ICML2022
Cited by in corpus (11)
- BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
- A Survey on Unsupervised Anomaly Detection Algorithms for Industrial Images
- DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
- Context Disentangling and Prototype Inheriting for Robust Visual Grounding
- VQA-based Robotic State Recognition Optimized with Genetic Algorithm
- Robotic Applications of Pre-Trained Vision-Language Models to Various Recognition Behaviors
- Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning
- Synthetic Boost: Leveraging Synthetic Data for Enhanced Vision-Language Segmentation in Echocardiography
- A request for clarity over the End of Sequence token in the Self-Critical Sequence Training
- Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
- OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification