PaLI: A Jointly-Scaled Multilingual Language-Image Model
arXiv:2209.06794
Abstract
Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision, language, and multimodal tasks, in many languages. To train PaLI, we make use of large pre-trained encoder-decoder language models and Vision Transformers (ViTs). This allows us to capitalize on their existing capabilities and leverage the substantial cost of training them. We find that joint scaling of the vision and language components is important. Since existing Transformers for language are much larger than their vision counterparts, we train a large, 4-billion parameter ViT (ViT-e) to quantify the benefits from even larger-capacity vision models. To train PaLI, we create a large multilingual mix of pretraining tasks, based on a new image-text training set containing 10B images and texts in over 100 languages. PaLI achieves state-of-the-art in multiple vision and language tasks (such as captioning, visual question-answering, scene-text understanding), while retaining a simple, modular, and scalable design.
ICLR 2023 (Notable-top-5%)
Cited by in corpus (11)
- Reproducible scaling laws for contrastive language-image learning
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Several categories of Large Language Models (LLMs): A Short Survey
- Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- Multimodal Neural Databases
- VLSlice: Interactive Vision-and-Language Slice Discovery
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval